METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Asana cuts browser agent model costs 76x

Asana ran a 144-run study with GPT-6 Astra in Codex and brought the per-run cost of the StackAI browser agent down to $0.47. Most of the savings came not from switching models but from caching the browsing history and removing screenshots in batches.

Asana cuts browser agent model costs 76x

Image: METAL

Summary

  • On October 9, OpenAI published a case study in which Asana optimized its browser agent around GPT-6.1 Sol, cutting model costs 76x and run time 5x.
  • GPT-6 Astra in Codex found the browsing history missing from the cache and ran a 144-run study across two history budgets, six policies and four models, three times each.
  • Optimization alone cut Model B's cost 29x, and with the larger history budget Sol returned the correct answer in all 18 runs.

Work management company Asana cut the model costs of its no-code browser agent 76x in tests. In a customer case study published on October 9, OpenAI said Asana ran experiments with GPT-6 Astra in Codex, OpenAI's coding agent, to optimize the browser agent's workflow around GPT-6.1 Sol. The estimated model cost per run fell to an average of $0.47, and run time fell to about four minutes, 76x cheaper and 5x faster than the original production setup.

The target was the browser agent of StackAI, an AI platform Asana acquired. On StackAI, customers build workflows that navigate websites, fill out forms and gather information without writing code. At Asana's scale, small inefficiencies in these workflows add up to large costs. StackAI CTO Frank Hidalgo, PhD, took on making the agent cheaper and faster, and handed the investigation, the testing of improvements and the comparison of results to GPT-6 Astra in Codex.

The first thing Astra found was a gap in caching. The agent cached its fixed instructions and tool definitions, but not the growing history of page text and screenshots it gathered while browsing, so every request resent that whole history at full price. The agent also dropped older screenshots and trimmed text at nearly every step, and each edit changed the history, so caching the history alone would not have helped. Losing those facts could also force the agent to revisit pages it had already read.

Hidalgo picked three of the fixes Astra proposed: extending caching to the browsing history, increasing the amount of text retained, and removing screenshots in batches rather than at every step. Astra began with quick tests to establish which variables mattered, then refactored code that had not been built for controlled experiments so that one frontend and backend could run many workflows with different settings in parallel.

The main study was 144 runs. History budgets of 120,000 and 480,000 characters and six caching and screenshot policies were each tested three times on four models: GPT-6.1 Sol and Models A, B and C from another frontier lab. Model A is a smaller model released in fall 2025 at half Sol's price, while Model B, used in production, and its updated version Model C are priced the same as Sol. The task was collecting six fields for each of 32 books from a public demo catalog, modeled on work Asana customers actually run on StackAI. The best policy let screenshots accumulate to 20 before cutting back to the most recent one. Between removals, earlier history stays unchanged for longer, which keeps the cache working.

Costs fell at every stage. Optimization alone brought Model B's per-run cost from at least $36.21 down to $1.24, a 29x reduction, and the optimized GPT-6.1 Sol workflow was a further 2.6x cheaper at $0.47. On Sol alone, with the larger history budget, the new caching and screenshot policy cut cost 4x, from $1.97 to $0.47 per run. Because 89% of the input came from cache at 5% of the uncached price, each call became about 3x cheaper. The original chart that METAL reviewed shows the optimized Model C also reaching $0.66 and 3.5 minutes. On run time alone, Model C was faster than Sol's 4.0 minutes, and most of the 76x gap came from fixing history management rather than switching models.

The history budget also decided whether the agent produced an answer at all. GPT-6.1 Sol produced answers in only three of 18 runs with the smaller budget, but returned the correct answer in all 18 with the larger one. OpenAI said every run in the optimized workflow completed the task and returned the correct answer. Every session's requests, data traces and results were recorded in Command, Asana's software delivery platform, and the findings went through tickets and pull requests into production.

아사나 브라우저 에이전트의 1회 비용과 실행 시간 비교 도표. 모델 B 개발 기준선 36.21달러·22.5분 이상에서 모델 B 최적화 1.24달러·4.7분, 모델 C 최적화 0.66달러·3.5분, GPT-6.1 Sol 최적화 0.47달러·4.0분으로 줄었다

"This would have taken me one to two months by hand," Hidalgo said. "With GPT-6 Astra in Codex, it took about a week: I'd set a /goal before going to bed and review the results in the morning." He said, "Cost used to limit which models we could offer customers for these workloads," explaining that making the agent more efficient lets the company offer a better, faster model while lowering operating costs. Arnab Bose, CPO at Asana, said, "An engineer set the direction, GPT-6 Astra ran the experiments, and the results went through Command to production," calling it what teams of humans and agents look like in practice.

METAL previously reported on OpenAI's launch of GPT-6.1 Sol. Asana has already shipped the changes to browser navigation in StackAI and is building tools that make similar experiments easier to repeat. Over time it plans to fold this testing into the platform's evaluations so customers and internal teams can compare cost, run time and answer quality when configuring their agents. Asana is also using Astra to test product features before release, with Astra navigating the platform, trying different inputs and reporting bugs to human QA reviewers. "Shipping speed is no longer the bottleneck; human attention is," Hidalgo said. The case shows that much of an agent's cost depends not on model pricing but on how its history is accumulated and trimmed, and that even the experiments to find that design are now being handed to agents.

Comments