Skip to content
Rush Commerce
Tools & Teardowns3 min read

Strands Decider 2B: AWS open-sources a 115ms decision model

AWS Strands Labs open-sourced Strands Decider 2B, a decision model that picks from fixed options with a confidence score. Use it to route tickets and gate tool calls.

AWS's Strands Labs released Strands Decider 2B on October 1, an open-source decision model. It does not write text. It picks one option from a list you give it and returns a confidence score with the pick. It runs on one consumer GPU or a MacBook. Most of the "AI" calls in a small-business agent are exactly this kind of job, and most teams pay a frontier model to do it.

What actually happened

Per the Strands announcement and the GitHub repo:

  • Base: Qwen3.5-2B with the language head removed and replaced by a "pointer head" that scores each option you pass in.
  • License: Apache 2.0. Weights are on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19. Install with pip install strands-decider.
  • Speed: about 115ms median on an RTX 3090 and about 153ms on an Apple M3 Pro for small tasks. Latency grows roughly linearly with task size.
  • Accuracy: 72.3% on JevBench (167 of 231 tasks) with a calibration error of 0.052. Strands says that puts it 3rd of 33 models in the 2B class on the benchmark's public set.
  • Use cases named: model routing, tool selection, evals, guardrails, memory and policy classification.

The CLI example in the repo is a support ticket: "My payouts have been failing for 3 days," with the choices billing, sales, retail. The SDK hooks into before_tool_call, so the model can approve or block a tool call before your agent runs it.

TechCrunch frames it as AWS's answer to TypeSafe's Jev, the decision model we covered earlier.

Why a local decision model matters for your business

Stop paying for prose you throw away. "Is this email a refund request, a sales lead or spam?" is a classification. If a frontier model answers it, you pay for a reasoning pass and then parse one word out of a paragraph. A 2B model on a box you already own answers in about a tenth of a second for the cost of electricity.

The confidence score is the useful part. A pick with a score lets you write a rule: above 0.9, act; below it, send to a person or to the big model. That is a hybrid agent, and it is how you cut model spend without cutting quality on the hard cases.

72% is not "set and forget." That benchmark number means roughly one wrong pick in four on JevBench tasks. Test it on 200 of your own labeled tickets before you let it act alone, and keep the threshold high at first.

Key takeaways

  • Strands Decider 2B picks from fixed options and returns a confidence score
  • Apache 2.0, built on Qwen3.5-2B, about 115ms median on an RTX 3090
  • 72.3% on JevBench with 0.052 calibration error, per the repo
  • Good fits: ticket routing, tool-call gating, model routing, policy checks
  • Gate actions on the confidence score and test on your own labeled data first

Paying frontier prices to sort your inbox? We build agents that send the easy decisions to a small local model and only the hard ones to the expensive one. Estimate the savings or see how we build.

Sources: Strands Agents, GitHub: strands-labs/strands-decider, TechCrunch.

  • #strands-decider
  • #aws
  • #decision-models
  • #ai-agents
  • #open-source
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.