Cognition's SWE-2 lands within a point of Fable 5.1
SWE-2 scores 50.0% on FrontierCode 1.1 at 64% lower cost than Fable 5.1, post-trained on an open 2.8T Chinese base model. What that means for your coding agent budget.
Cognition released SWE-2 on September 10, and the interesting number is not the benchmark — it is the gap between the benchmark and the bill. SWE-2 scores 50.0% on Cognition's FrontierCode 1.1 Main against Fable 5.1's 50.9%, a difference of nine tenths of a point, at 64% lower cost. If you are running a coding agent on frontier pricing out of habit, this is the release that makes you re-run the math.
What actually happened
Per Cognition's writeup, SWE-2 is post-trained from Kimi K3, a 2.8-trillion-parameter open model that had already been through heavy RL for agentic coding. Cognition's own reinforcement learning sits on top of that base. Published scores: 50.0% on FrontierCode 1.1 Main, 73.0 on DeepSWE 1.1, 92.8 on Terminal-Bench 2.1, and 27.3 on Terminal-Bench 4.
The mechanism worth reading is effort levels. Cognition trained medium, high, and max in a single RL run rather than shipping three separate models. Medium favors speed and costs 81% less on average than SWE-1.7 on the same FrontierCode tasks; high and max spend more compute on hard problems. One training run, three points on the cost curve.
Cognition has not published a per-token price. Every cost claim is relative: 64% below Fable 5.1 at a matched score, 81% below its own predecessor. SWE-2 shipped first in Devin Desktop and the Devin CLI, with Devin Web and Fusion following.
Why a cheaper coding agent matters for your business
Most teams pick one model and route everything through it. That is a rounding error on a hobby project and a real line item at ten engineers. Effort levels turn model choice into a per-task decision, which is where it belonged all along: a dependency bump and a payment-flow refactor should not cost the same.
Two caveats we would hold onto. First, FrontierCode is Cognition's own benchmark. A vendor grading itself against a competitor is useful signal and not an independent result — measure SWE-2 on your repo, on the kind of ticket you actually close, before you switch anything.
Second, no published token price means no contract. "64% cheaper than a competitor" is a ratio that moves when the competitor moves, and it is only reachable through Devin's surfaces today. That is a product decision, not a model you can drop behind your own router. Price the lock-in along with the savings.
The base model deserves a note too. SWE-2 sits on open weights from Kimi K3, a Chinese open release. Your coding agent now has a supply chain, and it runs through a lab you did not pick. Know that before someone in procurement asks.
Key takeaways
- SWE-2 shipped September 10 at 50.0% on FrontierCode 1.1 Main - within one point of Fable 5.1, at 64% lower cost
- Post-trained from Kimi K3, a 2.8T-parameter open model already RL-tuned for agentic coding
- Medium, high, and max effort levels came from one RL run; medium costs 81% less on average than SWE-1.7
- No published per-token price - every cost claim is relative, and access is through Devin surfaces only
- FrontierCode is Cognition's own benchmark; validate on your repo before switching
- Route by task difficulty instead of sending every ticket to one frontier model
Model pricing changes every few weeks. Your architecture should not. We build routing layers that let you move a workload to a cheaper model without touching your product code. Run the numbers on your automation spend, or bring us a model bill that stopped making sense.
Sources: Cognition: SWE-2.
- #cognition
- #coding-agents
- #devin
- #model-pricing
- #developer-tools
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
VS Code 1.137 puts your coding agent on a schedule
VS Code 1.137 ships Automations in preview: recurring agent tasks that run hourly, daily, or weekly. Here's what that changes about cost and review.
Read itSol-H3 generates video faster than it plays back
NVIDIA's Sol-H3 stack renders 5 seconds of 1344x768 video with stereo audio in 1.653 seconds on 8x B300. AI video generation just crossed real time. Here is what changes.
Read it