Strands Decider 2B: AWS open-sources a 115ms decision model
AWS Strands Labs open-sourced Strands Decider 2B, a decision model that picks from fixed options with a confidence score. Use it to route tickets and gate tool calls.
AWS's Strands Labs released Strands Decider 2B on October 1, an open-source decision model. It does not write text. It picks one option from a list you give it and returns a confidence score with the pick. It runs on one consumer GPU or a MacBook. Most of the "AI" calls in a small-business agent are exactly this kind of job, and most teams pay a frontier model to do it.
What actually happened
Per the Strands announcement and the GitHub repo:
- Base: Qwen3.5-2B with the language head removed and replaced by a "pointer head" that scores each option you pass in.
- License: Apache 2.0. Weights are on Hugging Face as
StrandsAgents/strands-decider-2B-hobson-v19. Install withpip install strands-decider. - Speed: about 115ms median on an RTX 3090 and about 153ms on an Apple M3 Pro for small tasks. Latency grows roughly linearly with task size.
- Accuracy: 72.3% on JevBench (167 of 231 tasks) with a calibration error of 0.052. Strands says that puts it 3rd of 33 models in the 2B class on the benchmark's public set.
- Use cases named: model routing, tool selection, evals, guardrails, memory and policy classification.
The CLI example in the repo is a support ticket: "My payouts have been failing for 3 days," with the choices billing, sales, retail. The SDK hooks into before_tool_call, so the model can approve or block a tool call before your agent runs it.
TechCrunch frames it as AWS's answer to TypeSafe's Jev, the decision model we covered earlier.
Why a local decision model matters for your business
Stop paying for prose you throw away. "Is this email a refund request, a sales lead or spam?" is a classification. If a frontier model answers it, you pay for a reasoning pass and then parse one word out of a paragraph. A 2B model on a box you already own answers in about a tenth of a second for the cost of electricity.
The confidence score is the useful part. A pick with a score lets you write a rule: above 0.9, act; below it, send to a person or to the big model. That is a hybrid agent, and it is how you cut model spend without cutting quality on the hard cases.
72% is not "set and forget." That benchmark number means roughly one wrong pick in four on JevBench tasks. Test it on 200 of your own labeled tickets before you let it act alone, and keep the threshold high at first.
Key takeaways
- Strands Decider 2B picks from fixed options and returns a confidence score
- Apache 2.0, built on Qwen3.5-2B, about 115ms median on an RTX 3090
- 72.3% on JevBench with 0.052 calibration error, per the repo
- Good fits: ticket routing, tool-call gating, model routing, policy checks
- Gate actions on the confidence score and test on your own labeled data first
Paying frontier prices to sort your inbox? We build agents that send the easy decisions to a small local model and only the hard ones to the expensive one. Estimate the savings or see how we build.
Sources: Strands Agents, GitHub: strands-labs/strands-decider, TechCrunch.
- #strands-decider
- #aws
- #decision-models
- #ai-agents
- #open-source
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
LlamaIndex Extract v2.5: cheapest tier beats old Agentic
LlamaIndex Extract v2.5 lifts document extraction accuracy on every tier at the same per-page price. Its Cost Effective tier now beats the old Agentic. Re-test your tier.
Read itDigitalOcean Agent Droplets: $50 a month, set the spend cap
DigitalOcean Agent Droplets bundle microVMs, open-model inference and 16,000 tools for $50 or $200 a month. Frontier models stay at list price. Set the cap first.
Read it