Asana's browser agent got 76x cheaper: the fix was caching
StackAI by Asana cut browser agent cost 76x on GPT-6.1 Sol. Most of the saving came from prompt caching and screenshot pruning, not the model swap.
OpenAI published a case study this week: Asana's browser agent now runs 76x cheaper and 5x faster on GPT-6.1 Sol. The headline sells the model. The data says something else. On a different model, the same workflow changes alone cut browser agent cost 29x. Prompt caching and how the agent keeps its screenshot history did most of the work. That is good news, because you can apply the same fix to the agent you already run.
What actually happened
The test was run by StackAI, the agent workflow platform Asana now offers as "StackAI by Asana." OpenAI's write-up and StackAI's own report give the details:
- The task: a browser agent collects six fields for each of 32 books from a public demo catalog. That is 192 facts, each scored against a reference answer.
- The study: 144 runs across GPT-6.1 Sol and three unnamed models (A, B and C). Six caching and history policies, two history budgets, three runs per condition. Then a 12-run follow-up.
- The model-agnostic win: on Model B, the best policy cut cost per run 29x and run time 4x against the original production setup.
- The Sol result: the best GPT-6.1 Sol setup averaged $0.47 in estimated model cost and about four minutes per run, which is 76x cheaper and 5x faster than the original setup on Model B. It read 89% of its input from cache. Our math: the model swap adds roughly 2.6x on top of the 29x from the workflow.
The fix had three parts. Cache the conversation history, not only the system prompt and tools. Keep up to 20 screenshots, then prune back to one, so about 19 calls in a row reuse the cache. And raise the history budget from 120,000 to 480,000 characters so the agent stops forgetting what it read.
StackAI is honest about the limits. Without pruning, caching the larger history cost more than no caching on three of four models. With three runs per condition, the study shows broad patterns, not small differences. OpenAI's Codex with GPT-6 Astra did much of the instrumentation and analysis.
Why prompt caching matters for your business
Browser agents are expensive because every step resends the whole history, screenshots included. A cache only reuses the longest unchanged prefix of a request. If your agent trims or rewrites old history on each call, the prefix breaks, and you pay full price every time. Cache reads cost 0.05x to 0.1x the normal input price on the models tested. That gap is where most of the 29x came from.
Before you change models, measure your cache hit rate. Then make the history append-only and prune in batches, not on every call.
Key takeaways
- StackAI by Asana cut browser agent cost 76x and run time 5x on GPT-6.1 Sol in a 144-run study
- Workflow changes alone gave 29x on another model; the model swap added roughly 2.6x
- The fix: cache the history, prune screenshots in batches of 20, raise the history budget
- Without batch pruning, caching the larger history cost more on three of four models
- Measure your cache hit rate before you pay to switch models
Running a browser or ops agent that costs more than it should? We audit agent loops for cache breaks, history bloat and wrong-tier model calls, then fix them in your stack. See how we build agents, or estimate your savings.
- #browser-agents
- #prompt-caching
- #gpt-6-1-sol
- #ai-costs
- #asana
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
GPT-6.1 Sol Ultrafast costs $12/$60: pay only where users wait
OpenAI's GPT-6.1 Sol Ultrafast is live in the API at $12/$60 per million tokens, 6x Standard. Use it only on the steps where a person is waiting.
Read itClaude submitted real web forms in tests: gate your agent's submit
Anthropic's unintended-actions report shows Claude submitting real forms and bypassing limits on live sites. Put a hard gate on every agent write action.
Read it