AWS open-sources Strands Decider 2B: a 2B model that never writes a word — and decides in 115 milliseconds
AWS's Strands Labs released Strands Decider 2B on October 1, 2026: an open-source ~2B-parameter model built on a Qwen3.5-2B torso with its text head swapped for a pointer that scores fixed options — it can't write a sentence, but it can gate an agent's tool call in about a tenth of a second.

Amazon Web Services introduced Strands Decider 2B on Thursday through its Strands Labs program: a roughly 2-billion-parameter open-source model that answers bounded questions with choices and probabilities in a single pass. It is built on the "torso" of Alibaba's Qwen3.5-2B — but it cannot generate text at all. The word-making head was cut off and replaced with a pointer that scores the options it is handed. The point is speed and cost: a decision in about a tenth of a second, cheap enough to sit in agent workflows where a full LLM call never could.
The text head, removed#
Strands Labs took the Qwen3.5-2B torso and removed the language-model head, replacing it with a pointer head of just over a million parameters that scores each answer option, plus a rank-16 LoRA adapter for the rest. The public release is v19 — an earlier slot-head design performed significantly worse, the team said. Because the model only answers from the options it is given, it runs fast and never goes off script, and AWS says each decision carries a reliability score that frontier LLM inference APIs do not expose. AWS quotes a median of under 100 milliseconds per decision; SQ Magazine's reporting cites about 115 milliseconds on a single Nvidia RTX 3090. On JevBench, the public benchmark for decision models, Decider 2B placed third of 33 models in its size class for accuracy and calibration combined — first of 30 once models just over 2 billion parameters are excluded — and answered 100% of the benchmark's easy tasks correctly. Calibration was measured with the Brier score, which checks whether stated confidence matches actual accuracy. Weights, training data, and training scripts are all downloadable from GitHub and Hugging Face; developers install it with pip install strands-decider and load the StrandsAgents/strands-decider-2B-hobson-v19 checkpoint.
A gate before the tool call#

AWS's own launch demo shows the intended job. A deliberately eager agent gets asked "What's the weather?" and guesses a city anyway. Before the get_weather tool runs, Decider answers two yes/no questions about grounded arguments and premature calls — and the agent goes back to ask which city the user meant. The check plugs into Strands' intervention system, where a before_tool_call handler can return Proceed, Deny, Confirm, or Guide. AWS picked the demo's questions, threshold, and policy by hand and calls it an illustration only. The team's pitch, quoted by SQ Magazine: "a decision this cheap can sit in a path where an LLM call never could."
The decision-model gold rush#

The project began as distinguished engineer Marc Brooker's homebrew take on TypeSafe's Jev, a model launched in September that chooses among predefined options instead of generating prose. Brooker's version briefly topped the JevBench ranking for its size, so AWS engineers cleaned it up and shipped it through Strands Labs, the company's group for agent tooling. TypeSafe named Jev after economist William Stanley Jevons, whose paradox holds that a cheaper resource can end up in higher demand. OpenAI announced a similar offering the same week, and TechCrunch notes dozens of researcher-built clones have appeared since Jev debuted. Brooker told TechCrunch the driver was AWS customers, whose agentic workflows "didn't always require the capability or cost of a fully featured LLM all the time." He doesn't expect frontier labs to dominate the category — an interesting niche model costs "hundreds or thousands of dollars" to build. TypeSafe CEO Diogo Almeida is unfazed: "I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart."
The honest caveats#
Bound the enthusiasm. The published latency chart measured v18, one version behind the release. The demo's questions, thresholds, and policy were hand-set and need testing on real traffic. A single parallel pass leaves Decider far behind reasoning models on complex problems — coding, chatbots, and document summaries are explicitly off the table. And the biggest open question: how its confidence scores hold up outside JevBench's public set.
Sources
- Strands Agents blog — "Introducing Strands Decider 2B: a small, open source, decision model," by Marc Brooker, Mike Chambers and Fabio Nonato de Paula (October 1, 2026) — the announcement
- TechCrunch — "Amazon releases its own Jev clone as decision models flood the web" (October 1, 2026) — Brooker and TypeSafe CEO Diogo Almeida quotes
- SQ Magazine — "Amazon Strands Decider 2B: Fast Open-Source Decision Model" (October 1, 2026) — architecture, benchmark numbers and launch demo