OpenAI Decisions API: 150ms picks on Luna, price not yet public
OpenAI's Decisions API returns a fixed-option answer in about 150ms on GPT-6 Luna. It is in limited preview with no published price. Here is how to test it.
OpenAI's new Decisions API does one job: you give it a question and a closed list of answers, and it gives back one of them. No paragraph, no parsing. OpenAI announced it at DevDay on September 29 as a limited preview. It is the latest in a run of decision-model launches, and the first from a frontier lab you probably already pay.
What actually happened
Per OpenAI's DevDay recap and The Decoder's coverage:
- How it works: you define questions with a finite set of predefined answers, send context as text or images, and get an answer back.
- Model: a version of GPT-6 Luna, OpenAI's low-cost model. Luna lists at $0.10 input and $0.50 output per million tokens in the regular API.
- Speed: about 150 milliseconds, which The Decoder reports as roughly ten times faster than asking Luna the same thing through the normal API.
- Named use cases: classify content, route requests, choose an agent's next action.
- Availability: limited preview, with broad release "in the coming days."
What OpenAI has not said: the price. There is no public rate card yet, and no word on whether it bills per call, per token or per question. Do not assume it costs the same as a plain Luna call.
It lands in a crowded week. AWS open-sourced Strands Decider 2B, and TypeSafe's Jev started the category.
Why a decisions API matters for your business
Most of your agent's calls are decisions. "Is this a refund, a lead or spam?" "Does this order need a human?" "Which tool runs next?" Each one is a pick from a short list. Today most teams send those to a chat model and regex one word out of the reply. A decision endpoint removes the parsing step and the latency that goes with it.
150ms changes where you can put AI. At that speed a check fits inside a checkout flow or a live chat turn, not just a background job.
Pricing decides the winner, and it is blank. A hosted API means no GPU to run. A local model like Strands Decider means a fixed cost and no data leaving your box. Until OpenAI publishes a price, you cannot do that math. Build your routing behind one interface so you can swap the backend when the numbers land.
Measure accuracy on your own data. Pull 200 labeled tickets or orders, run them through the preview and your current setup, and compare. Speed is worth nothing if one pick in five is wrong.
Key takeaways
- Decisions API returns one answer from a fixed list, given text or image context
- Runs on a version of GPT-6 Luna at about 150ms
- Limited preview now; broad release promised in days
- No published price yet, so do not budget on Luna's token rates
- Put decisions behind one interface and benchmark against local models
Your agent asks a chat model to pick from a list? We rebuild those calls as typed decisions with a fallback to a person, and keep the backend swappable. See how we build or estimate the savings.
Sources: OpenAI DevDay 2026 recap, The Decoder.
- #openai
- #decisions-api
- #gpt-6-luna
- #decision-models
- #ai-agents
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Dynatrace closes $915M Arize deal: keep AI traces on OpenTelemetry
Dynatrace closed its $915M Arize acquisition. Phoenix stays open source and OpenInference now lives in OpenTelemetry — so instrument your AI once, on the standard.
Read itRestate's $20M Series A: durable execution, priced per action
Restate raised $20M to make durable execution cheap for AI agents. Its own pitch: managed runtimes charge $25–50 per million actions. Price your steps first.
Read it