OpenAI Jalapeño benchmarks: your token price has a new floor
OpenAI published the first Jalapeño inference benchmarks at Hot Chips 2026. Here is what a custom OpenAI inference chip does to what you pay per token.
OpenAI put real numbers behind its custom silicon this week. At Hot Chips 2026 the company published the first benchmark results for Jalapeño, an inference chip co-designed with Broadcom and built for serving language models rather than training them. The numbers are good. The part that reaches your invoice is simpler: the company that sells you tokens is now building the machine that makes them.
What actually happened
Per TechCrunch, OpenAI presented Jalapeño on Tuesday with results measured on SemiAnalysis's InferenceX suite. Richard Ho, who runs hardware at OpenAI, framed the goal as more AI work per unit of power and faster responses at the same time — the two things that usually trade against each other.
SemiAnalysis's own writeup has the specifics. On DeepSeek R1 at concurrency 1, Jalapeño cleared 700 tokens per second per user. On Kimi K2.5 it landed near 700 tok/s/user against roughly 100 for the next best part measured. On GPT-OSS, throughput at matched interactivity approached double GB200's best point. Disclosed hardware: 700W TDP per compute die, 15.4TB/s of HBM4 bandwidth, 13.4 PFLOPs MXFP4 on the B0 stepping, built on N3P. On cost per output token, SemiAnalysis puts Jalapeño roughly level with Vera Rubin — and Jalapeño got there without speculative decoding turned on.
Deployment is limited and internal through the end of 2026, with a larger ramp in 2027. Nobody is buying one.
Why a cheaper inference chip matters for your business
This is a cost-floor story, not a chip story. You will never touch Jalapeño. You will feel it in per-token pricing eighteen months from now, the same way you felt every previous efficiency jump. Throughput per megawatt is the number that eventually sets what an API call costs.
Do not sign a long AI contract into a falling market. Anyone quoting you a three-year committed spend on inference is asking you to prepay at 2026 prices. Keep terms short. Keep the option to move.
Vertical integration cuts both ways. A model vendor that owns its silicon can lower your price. It can also make its cheapest path the only good one. Design so that swapping a model provider is a config change, not a rewrite: one internal interface for model calls, prompts and tools in your repo, outputs stored in your schema.
Latency at concurrency 1 is the number that matters for agents. A support agent or a voice bot is one user waiting. 700 tokens per second per user is the difference between a conversation and a hold message. If you are benchmarking vendors, test single-user latency, not batch throughput.
Key takeaways
- OpenAI published its first Jalapeño benchmarks at Hot Chips 2026 on August 25; the chip is co-designed with Broadcom
- SemiAnalysis measured over 700 tok/s/user at concurrency 1 on DeepSeek R1 and near-parity with Vera Rubin on cost per output token
- Limited internal deployment through end of 2026, wider ramp in 2027 — this is not a product you can buy
- Keep AI contracts short. Inference efficiency is still improving and prices follow it down
- Benchmark vendors on single-user latency, not batch throughput, if you run agents
Cheap tokens only help if you can switch to them. We build AI systems with one swappable model layer, your prompts in your repo, and your data in your database — so a price drop is a config change, not a rebuild. See how we build it.
Sources: TechCrunch, SemiAnalysis.
- #openai
- #inference
- #jalapeno
- #broadcom
- #ai-costs
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Ringg raises $10M: voice AI is really an outcome layer
Peak XV put $10M into Ringg, a voice AI startup running 20M call attempts a month. The lesson: buy the workflow engine, not the channel.
Read itThomson Reuters built its own model for $40M — data is the asset
Thomson Reuters trained a proprietary LLM on Westlaw and Reuters content for $40M and owns it outright. The lesson for operators isn't build-your-own — it's what you own.
Read it