OpenAI tests outcome-based pricing after an 80% price cut
OpenAI's CFO says an 80% cut to Luna drove 10x usage, and the company is testing outcome-based pricing. Re-run your AI cost math before the meter changes.
OpenAI CFO Sarah Friar spent her time at Goldman Sachs' Communacopia + Technology Conference on two numbers that matter more to your budget than any benchmark. The company cut the price of its low-cost Luna model by 80% and saw usage jump roughly tenfold. And it is experimenting with outcome-based pricing instead of charging per token. Those two facts point in opposite directions, and the gap between them is your planning window.
What actually happened
Speaking in San Francisco on September 8, Friar said enterprise revenue grew 32% from June to July against 20% growth in overall annualized revenue over the same stretch, and that the enterprise and consumer businesses reached a roughly even split by mid-year — ahead of OpenAI's own year-end target, according to Reuters.
The vertical push is the headline: chip design, life sciences, financial services. OpenAI used its own models to design its Jalapeno chip and reached tape-out in nine months. Friar's competitive claim was explicitly about price, not capability — deploying Luna against Z.ai's GLM 5.3 on a cloud layer, she said, "we are cheaper." Codex, meanwhile, has grown to 25 million users from roughly 100,000 at the start of the year.
Then the part nobody quoted in a headline: OpenAI is testing pricing tied to outcomes rather than usage, because buyers keep asking for a clearer return.
Why it matters for your business
Start with the 10x. AI pricing elasticity is now measured, not theoretical. An 80% cut produced roughly ten times the volume, which means the price cut was not generosity — it was a demand experiment that paid. The operator lesson is unglamorous and immediate: any automation you priced out in the last twelve months deserves a fresh spreadsheet. The document-extraction job that cost $400/month at old rates might cost $80 now. That's the difference between "someday" and "this quarter."
The outcome-pricing experiment is the one to watch warily. Per-token billing is ugly, but it is the only AI pricing model you can actually forecast: you know your call volume, you know your context size, you can cap it. Outcome pricing moves the meter to something the vendor defines — a resolved ticket, a completed task, a closed matter — and it prices your automation against the value it creates rather than the compute it burns. That is a fine deal when the automation is marginal and a terrible one when it works. The better it works, the more it costs.
So do two things while tokens are still metered and cheap. Re-run the math on the automations you shelved. And keep the model layer swappable, so that when a vendor changes how it charges, moving is a config edit rather than a negotiation you enter with no leverage.
Key takeaways
- An 80% price cut on OpenAI's Luna model drove roughly 10x usage — elasticity in this market is now a measured fact
- OpenAI enterprise revenue grew 32% June to July vs 20% for overall annualized revenue; enterprise and consumer are now roughly even
- Codex went from ~100,000 users at the start of 2026 to 25 million
- OpenAI is testing outcome-based pricing — that shifts the meter from compute you can forecast to value the vendor defines
- Two moves now: re-cost the automations you shelved, and keep model choice in config so a pricing change isn't a lock-in
You priced that automation at last year's token rates. The floor moved. Put your real volumes into our ROI calculator and see what the job costs today, or send us the workflow you gave up on and we'll tell you straight whether the numbers work now.
Sources: Reuters, via Investing.com.
- #ai-pricing
- #openai
- #token-costs
- #vendor-risk
- #automation-roi
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Suno v6 retrained on licensed music. Check your lineage
Suno rebuilt v6 from scratch on licensed Warner, BMG and Believe catalogs and retired its old models the same day. Training-data provenance is now a product feature.
Read itMercury 2.5 hits 1,107 tokens/sec: price the fast path
Inception's Mercury 2.5 diffusion LLM runs 1,107 tokens per second at $0.20/$0.75 per million. Latency is now a routing decision, not a model you are stuck with.
Read it