{"id":5932,"date":"2026-10-10T01:20:28","date_gmt":"2026-10-10T04:20:28","guid":{"rendered":"https:\/\/tucumandevelopers.com\/index.php\/2026\/10\/10\/asana-cuts-model-costs-76x-in-browser-tests-with-gpt-6-1-sol\/"},"modified":"2026-10-10T01:20:28","modified_gmt":"2026-10-10T04:20:28","slug":"asana-cuts-model-costs-76x-in-browser-tests-with-gpt-6-1-sol","status":"publish","type":"post","link":"https:\/\/tucumandevelopers.com\/index.php\/2026\/10\/10\/asana-cuts-model-costs-76x-in-browser-tests-with-gpt-6-1-sol\/","title":{"rendered":"Asana cuts model costs 76x in browser tests with GPT-6.1 Sol"},"content":{"rendered":"<div>\n<div>\n<div id=\"identifying-browser-agent-inefficiencies-with-gpt-6-astra\">\n<p><h2><span>Identifying browser-agent inefficiencies with GPT\u20116 Astra<\/span><\/h2>\n<\/p>\n<\/div>\n<p><span>To move quickly, Hidalgo started by using GPT\u20116 Astra in Codex to map the codebase and explain how the agent built each model request. GPT\u20116 Astra discovered that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered, so every request resent that history at full price. <\/span><\/p>\n<p><span>The agent also dropped older screenshots and trimmed text at nearly every step. Each edit altered the history, so caching the history alone would not have helped, and losing those facts could require the agent to revisit pages it had already read.<\/span><\/p>\n<div id=\"from-an-estimated-two-months-of-research-to-one-week-with-gpt-6-astra\">\n<p><h2><span>From an estimated two months of research to one week with GPT\u20116 Astra<\/span><\/h2>\n<\/p>\n<\/div>\n<p><span>Hidalgo reviewed GPT\u20116 Astra\u2019s proposed fixes and selected three to test:<\/span><\/p>\n<div>\n<ul>\n<li>\n<p><span>Extending caching to the agent\u2019s browsing history<\/span><\/p>\n<\/li>\n<li>\n<p><span>Increasing the amount of text it could retain<\/span><\/p>\n<\/li>\n<li>\n<p><span>Removing screenshots in batches rather than at every step<\/span><\/p>\n<\/li>\n<\/ul>\n<\/div>\n<p><span>GPT\u20116 Astra began with quick tests to establish which variables mattered. Because the code was not designed for controlled experiments, it then refactored the code so one frontend and backend could support many workflows in parallel, each with its own settings.<\/span><\/p>\n<p><span>Astra conducted the full study: history budgets of 120,000 and 480,000 characters and six caching and screenshot policies, each tested three times on each of the four models (see the table below). The best-performing policy allowed screenshots to accumulate to 20 before cutting back to the most recent one. This kept earlier history unchanged for longer stretches between removals. Combined with the larger history budget, it became the optimized workflow. Each configuration performed the same task: collecting six fields for each of 32 books from a public demo catalog, representative of what some Asana customers run in StackAI.<\/span><\/p>\n<div tabindex=\"0\">\n<table>\n<tbody>\n<tr>\n<th>\n<p><b>Model<\/b><\/p>\n<\/th>\n<th>\n<p><b>Description<\/b><\/p>\n<\/th>\n<th>\n<p><b>Price<\/b><\/p>\n<\/th>\n<\/tr>\n<tr>\n<td>\n<p><b>Model A<\/b><\/p>\n<\/td>\n<td>\n<p>A smaller, less expensive model from another frontier lab, released Fall 2025<\/p>\n<\/td>\n<td>\n<p>Half the price of <!-- -->GPT\u20116.1<!-- --> Sol<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><b>Model B<\/b><\/p>\n<\/td>\n<td>\n<p>The model originally used in production, from the same lab as Model A, released Summer 2026<\/p>\n<\/td>\n<td>\n<p>Same price as <!-- -->GPT\u20116.1<!-- --> Sol<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><b>Model C<\/b><\/p>\n<\/td>\n<td>\n<p>An updated version of Model B, released Fall 2026<\/p>\n<\/td>\n<td>\n<p>Same price as <!-- -->GPT\u20116.1<!-- --> Sol<\/p>\n<\/td>\n<\/tr>\n<tr>\n<td>\n<p><b>GPT\u20116.1<!-- --> Sol<\/b><\/p>\n<\/td>\n<td>\n<p>OpenAI\u2019s model<\/p>\n<\/td>\n<td><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><span>GPT\u20116 Astra ran the workflows and examined the requests, usage records and outputs, and separate model sessions reviewed the work. Every session\u2019s requests, data traces and results were recorded in <\/span><a href=\"https:\/\/asana.com\/product\/command\" target=\"_blank\" rel=\"noopener noreferrer\"><u><span>Command<\/span><\/u>\u2060<span>(opens in a new window)<\/span><\/a><span>, Asana\u2019s software delivery platform, so the team could review the complete study afterward. From Command, the findings were turned into tickets, then pull requests, and the changes went to production.<\/span><\/p>\n<div>\n<blockquote><p>\u201cThis would have taken me one to two months by hand. With GPT-6 Astra in Codex, it took about a week: I\u2019d set a \/goal before going to bed and review the results in the morning.\u201d<\/p><\/blockquote>\n<p>\u2014Frank Hidalgo, PhD, StackAI CTO at Asana<\/p>\n<\/div>\n<div id=\"bringing-model-cost-below-dollar050-per-run\">\n<p><h2><span>Bringing model cost below $0.50 per run<\/span><\/h2>\n<\/p>\n<\/div>\n<p><span>For Model B, the optimization reduced estimated model cost from at least $36.21 (some original runs reached the step limit before they finished) to $1.24 per run, a reduction of 29x. The optimized workflow on GPT\u20116.1 Sol was 2.6x cheaper still, at $0.47. Every run in the optimized workflow completed the task and returned the correct answer.<\/span><\/p>\n<div>\n<figure>\n<div data-tabs-autoplay-content=\"true\"><figcaption>\n<div>\n<p><span>Means of 3 runs. \u2265: the baseline includes capped runs, so its mean is a lower bound.<\/span><\/p>\n<p><span>The two right folds compare against Model B optimized. Model B ran in phase 1, Model C and Sol 6.1 in phase 2 of the same study (dotted line).<\/span><\/p>\n<\/div>\n<\/figcaption><\/div>\n<\/figure>\n<\/div>\n<p><span>On GPT\u20116.1 Sol alone, with the larger history budget, the new caching and screenshot policy reduced cost 4x, from $1.97 to $0.47 per run. Each call was about 3x cheaper, because 89% of the input came from cache at 5% of the uncached price. Runs also became faster: at least 22.5 minutes on the original setup on Model B, roughly four minutes with the optimized workflow on GPT\u20116.1 Sol.<\/span><\/p>\n<div>\n<figure>\n<div data-tabs-autoplay-content=\"true\"><figcaption>\n<div>\n<p><span>Mean of 3 runs, SD whiskers. \u2265: mean includes a capped or unfinished run, so the true value is at least this large.<\/span><\/p>\n<p><span>Bars use the blue theme. Read caching effects against the 480k larger-budget bar.<\/span><\/p>\n<p><span>Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.<\/span><\/p>\n<\/div>\n<\/figcaption><\/div>\n<\/figure>\n<\/div>\n<div>\n<figure>\n<div data-tabs-autoplay-content=\"true\"><figcaption>\n<div>\n<p><span>Mean of 3 runs, SD whiskers. \u2265: mean includes a capped or unfinished run, so the true value is at least this large.<\/span><\/p>\n<p><span>Bars use the blue theme. Read caching effects against the 480k larger-budget bar.<\/span><\/p>\n<p><span>Run markers and SD whiskers are approximate reconstructions from the source image; underlying run values and standard deviations were not available.<\/span><\/p>\n<\/div>\n<\/figcaption><\/div>\n<\/figure>\n<\/div>\n<p><span>The investigation also showed how history management affected whether the agent produced an answer at all. Giving GPT\u20116.1 Sol more room to retain its browsing history increased the number of runs that produced an answer from three of 18 with the smaller history budget to all 18 with the larger budget, each with the correct answer. For Hidalgo, the business value is giving customers access to faster, more capable models while keeping operating costs sustainable.<\/span><\/p>\n<div>\n<blockquote><p>\u201cCost used to limit which models we could offer customers for these workloads. By making the agent more efficient, we can give customers a better, faster model while lowering our operating costs.\u201d<\/p><\/blockquote>\n<p>\u2014Frank Hidalgo, PhD, StackAI CTO at Asana<\/p>\n<\/div>\n<div id=\"scaling-experimentation-and-product-testing\">\n<p><h2><span>Scaling experimentation and product testing<\/span><\/h2>\n<\/p>\n<\/div>\n<p><span>Asana has released the changes to browser navigation in StackAI and is developing tools to make similar experiments easier to repeat. Over time, the team plans to incorporate this testing into the platform\u2019s evaluations, so customers and internal teams can compare cost, runtime, and answer quality when configuring their agents.<\/span><\/p>\n<div>\n<blockquote><p>\u201cShipping speed is no longer the bottleneck; human attention is. We\u2019re close to a world where every engineer is a PM leading a fleet of agents.\u201d<\/p><\/blockquote>\n<p>\u2014Frank Hidalgo, PhD, StackAI CTO at Asana<\/p>\n<\/div>\n<p><span>Asana is now using GPT\u20116 Astra in Codex to test product features before release: Astra navigates the platform, tries different inputs and reports bugs for human QA reviewers. Hidalgo views this as the foundation for a new software development lifecycle, with many cloud agent sessions testing features in parallel.<\/span><\/p>\n<\/div>\n<\/div>\n<\/div>\n<\/div>\n<p>Fuente: <a href=\"https:\/\/openai.com\/index\/asana-browser-agent\">Art\u00edculo original<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Identifying browser-agent inefficiencies with GPT\u20116 Astra To move quickly, Hidalgo started by using GPT\u20116 Astra in Codex to map the codebase and explain how the agent built each model request. GPT\u20116 Astra discovered that the agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":5931,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":true,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"webixso_pending_account_ids":""},"categories":[43],"tags":[],"class_list":["post-5932","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-openai"],"jetpack_publicize_connections":[],"_links":{"self":[{"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/posts\/5932","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/comments?post=5932"}],"version-history":[{"count":0,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/posts\/5932\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/media\/5931"}],"wp:attachment":[{"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/media?parent=5932"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/categories?post=5932"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tucumandevelopers.com\/index.php\/wp-json\/wp\/v2\/tags?post=5932"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}