Ember-1 hits Kimi K3 quality with 40% fewer tokens
Fireworks trained a model to stop over-thinking: same coding benchmarks as Kimi K3, 71% fewer reasoning tokens. Cost per task, not cost per token.
The cheapest model is not the one with the lowest price per token. It is the one that finishes the job in fewer of them. Fireworks Research released Ember-1, a model built on Kimi K3 and trained for exactly that: same output quality, roughly 40% fewer tokens to get there. If you run coding agents, this is the axis that actually moves your bill.
What actually happened
Ember-1 landed on September 23 as a research preview on Fireworks Serverless, with two-week access. The pitch is narrow and specific: deliver Kimi K3's quality while cutting the reasoning tokens that agentic workloads burn between the prompt and the answer.
The benchmark table backs most of it up. On Terminal Bench 2.1, Ember-1 scores 82.0% against K3's best of 80.9%. On DeepSWE 1.1 it posts 75.2% versus 66.4%. On SWE-bench Verified it comes in slightly behind — 92.2% against K3's 93.2%. Call that a wash on capability, which is the honest read.
The token numbers are the story. Fireworks reports roughly 35% fewer tokens per task in live A/B testing, with total generated tokens down 39% and reasoning tokens specifically down 71.3%. The model, in their framing, "learned to cut unnecessary reasoning" — they say they developed new training algorithms to shorten reasoning without losing accuracy, and leave the technical details there.
On cost-per-task they claim Ember-1 beats GPT-6 Astra and Claude Opus 5, and they set a Pareto frontier on Bedside Bench, Doximity's physician-validated benchmark of 500 clinical cases across 10 categories.
Two caveats worth naming. Fireworks publishes Kimi K3's public API pricing ($3/M uncached input, $0.30/M cached, $15/M output) as the comparison baseline but does not state separate Ember-1 pricing in the announcement — so "40% fewer tokens" is a token claim, not a dollar claim, until you see your invoice. And the blog does not say whether the weights are open. A two-week research preview is a trial, not a dependency.
Why token efficiency matters more than token price for your business
Agent loops multiply everything. A single chat turn at 40% fewer tokens saves you pennies. An agent that runs 60 tool calls per task, 400 times a day, compounds that into real money. Reasoning tokens are the part of the bill nobody budgets for because nobody sees them in the UI.
Cost per task is the only metric that survives contact with a CFO. Price per million tokens is a vendor metric. Your finance team wants to know what it costs to triage one support ticket or open one pull request. A model that is twice the price per token and uses a third of them is cheaper, and no pricing page will tell you that.
A 1% benchmark loss can be the right trade — or fatal. Ember-1 gives up a point on SWE-bench Verified. For an internal refactor agent with human review, that is noise. For an unattended pipeline that ships without a second pair of eyes, that 1% is your incident rate. Decide which one you are running before you swap.
Benchmark on your own workload or do not swap. Every number above came from the vendor. Run your actual tasks through both models, log tokens and wall-clock per completed task, and compare. We have seen "efficient" models lose their advantage on domain-specific prompts because they over-reason on unfamiliar context.
Keep the swap cheap. This is why we route model calls through a single interface in every system we build. A two-week research preview is exactly the kind of thing you want to test in an afternoon and abandon without a migration.
Key takeaways
- Ember-1 from Fireworks Research is built on Kimi K3 and trained to use roughly 40% fewer tokens for the same quality
- Benchmarks: 82.0% on Terminal Bench 2.1 (K3: 80.9%), 75.2% on DeepSWE 1.1 (K3: 66.4%), 92.2% on SWE-bench Verified (K3: 93.2%)
- Live A/B tests show ~35% fewer tokens per task, 39% fewer total tokens, and a 71.3% cut in reasoning tokens
- Fireworks claims better cost-per-task than GPT-6 Astra and Claude Opus 5, and a Pareto frontier on Doximity's 500-case Bedside Bench
- Released September 23 as a two-week research preview on Fireworks Serverless — no separate Ember-1 price published, no stated weight license
- Reasoning-token efficiency compounds in agent loops, where a single task can span dozens of model calls
- Measure cost per completed task on your own workload; vendor benchmarks do not predict your prompts
Can you swap the model behind your agents in an afternoon? We build AI systems with the model behind an interface, token and cost-per-task logging on by default, so testing a new one is a config change and not a rewrite. Estimate what your agent loop costs today, or talk to us about the stack you are on.
Sources: Fireworks Research: Ember-1.
- #llm
- #token-costs
- #coding-agents
- #fireworks
- #benchmarks
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Gortex: a codebase graph your coding agent queries over MCP
Gortex indexes your repo into a knowledge graph and serves it to coding agents over MCP, so they ask for three functions instead of grepping whole files.
Read itClaude Code can now block new model releases by policy
Claude Code 2.1.283 adds deniedModels and availableModelsMatch managed settings, so a new model release can't enter your workflow until you approve it.
Read it