GPT-6.1 Sol Ultrafast costs $12/$60: pay only where users wait
OpenAI's GPT-6.1 Sol Ultrafast is live in the API at $12/$60 per million tokens, 6x Standard. Use it only on the steps where a person is waiting.
OpenAI is rolling out GPT-6.1 Sol Ultrafast in the API, Codex and ChatGPT Work, and the price is now on the page: $12 per million input tokens and $60 per million output. That is six times what the same model costs on the Standard tier. OpenAI says you get "near-Astra intelligence at up to 8x faster speeds." Speed is now a line item. The question for your agents is not "is it faster?" It is "which step is worth 6x?"
What actually happened with GPT-6.1 Sol Ultrafast
OpenAI announced the rollout on its developer forum this week. The details from the Ultrafast docs and the pricing page:
- Price per million tokens (up to 272K input):
gpt-6.1-solcosts $2 / $10 on Standard, $4 / $20 on Fast, and $12 / $60 on Ultrafast. Cached input is $0.60 on Ultrafast. - Astra is worse. GPT-6 Astra Ultrafast lists at $60 input and $300 output. OpenAI's own pitch is that Sol Ultrafast costs "just 1.2x" what Astra costs on Standard.
- How to turn it on: set
service_tiertoultrafaston each request. OpenAI strongly recommends WebSockets for agent loops with many short tool calls, because HTTP overhead eats the gain. - Separate rate limits: Sol Ultrafast starts at 1M tokens per minute on the Build tier, 4M on Launch and 40M on Grow. These limits are separate from Standard and Fast.
- Data residency: Sol Ultrafast supports US and EU residency. Astra Ultrafast is US only.
- Seats: in Codex and ChatGPT Work, Ultrafast is on Pro 500, eligible usage-based Enterprise and credit-based Edu plans. Enterprise admins must enable it.
The "up to 8x" speed figure is OpenAI's claim. The docs give no tokens-per-second number for the API.
Why it matters for your business
We said this when Ultrafast first appeared on GPT-5.6 Sol: faster tokens do not make your database, your payment API or your retries faster. If model generation is 30% of your agent's wall clock, an 8x model speedup saves about a quarter of the total time. You pay 6x on every token for it.
So split the work by who is waiting:
- A person is waiting: a live chat answer, a support agent on a call, an operator stuck on an outage. Ultrafast can pay here.
- Nobody is waiting: overnight catalog enrichment, invoice extraction, report drafts. Use Standard, or the Batch API. Speed you cannot see is money you lose.
- Cache first. A cached input token on Ultrafast costs $0.60, against $12 uncached. A stable prompt prefix cuts the premium more than any other change.
Key takeaways
- GPT-6.1 Sol Ultrafast is $12 input / $60 output per million tokens, 6x the Standard price of $2 / $10
- OpenAI claims up to 8x faster output; the API docs publish no tokens-per-second figure
- Enable it per request with
service_tier: "ultrafast", preferably over WebSockets - Ultrafast has its own rate limits and supports US and EU data residency for Sol
- Route only user-facing, latency-bound steps to Ultrafast; keep background work on Standard
Not sure where your agent spends its time? We instrument agent loops end to end, then route each step to the cheapest tier that meets its deadline. Run the numbers, or ask us to profile your workflow.
Sources: OpenAI Developer Community, OpenAI Ultrafast docs, OpenAI API pricing.
- #openai
- #gpt-6-1-sol
- #ultrafast
- #api-pricing
- #ai-agents
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Asana's browser agent got 76x cheaper: the fix was caching
StackAI by Asana cut browser agent cost 76x on GPT-6.1 Sol. Most of the saving came from prompt caching and screenshot pruning, not the model swap.
Read itClaude submitted real web forms in tests: gate your agent's submit
Anthropic's unintended-actions report shows Claude submitting real forms and bypassing limits on live sites. Put a hard gate on every agent write action.
Read it