What actually happened
In a case study published on 9 October 2026, OpenAI describes how Asana used GPT-6 Astra inside Codex to investigate and optimize a browser agent built with StackAI. The improved workflow running on GPT-6.1 Sol averaged $0.47 in estimated model costs and about four minutes per run. OpenAI reports a 76-fold model-cost reduction and about a fivefold speed improvement against the original production configuration.
The work involved a controlled task collecting six fields for each of 32 books from a public demo catalog. The engineering changes let the agent retain more useful browser history, cache more of its input and remove old screenshots in batches rather than repeatedly changing the request prefix. OpenAI says about 89% of the optimized agent’s input was served from cache. The comparison includes changes to workflow settings and model choice, so the 76x figure is not a pure model-price comparison.
Why a small team cares
Imagine a small distributor whose staff copy order details from a supplier portal into an internal spreadsheet. There may be no reliable API, and a browser agent looks tempting until the agent repeats navigation steps, forgets what it saw and burns tokens on screenshots. The total cost is the model bill plus retries, operator review and any bad records that must be corrected.
Asana’s example suggests checking the request construction before buying a cheaper model. Stable prompts and reusable context can reduce billing, while enough retained history can prevent repeated browsing. For a small operation, the benchmark should use a real weekly task with representative pages, interruptions and awkward inputs. Divide all run costs by the number of correctly completed tasks, then decide whether unattended use is justified.
Hype vs useful
The $0.47 is an estimated model cost in Asana’s specified test, not a guaranteed end-to-end operating cost or vendor price per job. It excludes other infrastructure and human costs. The headline improvement also starts from an expensive baseline, and the original setup had runs that hit a step limit. Its striking percentage should not be transferred to an unrelated portal.
The repeatable finding is narrower and more useful: browser-agent economics depend heavily on context design, caching policy and completion accuracy. Preserve the evidence and compare against a working deterministic integration whenever one exists.
