
Most agent decisions don't need an essay — they need a choice. Which team owns this ticket? Should the agent call the refund tool? Is this message urgent? Today those questions go through a full-size LLM with JSON parsing on the other side: slow, expensive, and brittle the moment the model gets chatty. Strands Decider (strands-labs/strands-decider, 503 GitHub stars, Apache-2.0) is a 1.9B-parameter decision model built for exactly this: it picks between options, answers yes/no, or rates on a scale — no text generation at all — with a calibrated confidence on every answer. Median answer time on an RTX 3090 is 115 milliseconds.
It comes from Strands Labs, the experimental arm of Strands Agents, and rides the current wave of "system one" decision models: small models that answer the fast, rote questions so your big LLM can spend its tokens on the hard ones. Under the hood it's a Qwen3.5-2B-Base torso with the language-modeling head thrown away, replaced by a ~1-million-parameter pointer head that scores each option against the model's answer position. One forward pass, no decoding loop; the torso is adapted with a rank-16 LoRA adapter and the head runs in fp32. Because the head carries no per-option parameters, your label sets are defined by the request — never baked into the weights.
pip install strands-decider is the whole install. The 1.9B checkpoint runs on an Apple-silicon Mac (MPS, or --device mlx for 1.4–1.6× faster than MPS), on CPU, or on a CUDA GPU — the 115 ms figure is the RTX 3090 reference.Install the CLI and triage a support ticket — a classic choice question, straight from the project's README:
pip install strands-decider
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 \
--state "Help! My payouts have been failing for 3 days! " \
--choice "Which team should handle this?=billing,sales,retail"
Example output from the project:
choice_0 -> billing (confidence 0.835)
billing 0.890
sales 0.056
retail 0.054
You get the pick, a confidence number, and a distribution over every option. That confidence is the feature: on short classification tasks it has never seen, answers at a confidence of 0.9 or more are right about 95% of the time (evaluation/results.md in the repo). Below that, the README says to confirm or ask a person. In an agent, this maps directly to a policy: act above your threshold, escalate below it.
The second question type is a yes/no — called noul, measuring how strongly the model leans "yes":
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 \
--state "Help! My payouts have been failing for 3 days! " \
--noul "Does this convey urgency?"
noul_0 noul = 0.875
Closer to 1 leans more toward yes. This is your guardrail primitive: policy classification, grounding checks, "should a human see this before it goes out."
The third type rates things on a scale — useful for evaluations and prioritization:
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 \
--state "Help! My payouts have been failing for 3 days! " \
--score "How frustrated is the writer?=calm,frustrated,depressed"
score_0 score = 1.07 (confidence 0.602)
0: calm 0.140
1: frustrated 0.648
2: depressed 0.212
The state is read once and each question adds only its own tokens, so combining questions is cheap:
strands-decider ask StrandsAgents/strands-decider-2B-hobson-v21 \
--state "Help! My payouts have been failing for 3 days! " \
--choice "Which team should handle this?=billing,sales,retail" \
--noul "Does this convey urgency?" \
--score "How frustrated is the writer?=calm,frustrated,depressed"
You now have routing, urgency, and frustration in a single batched call — the three things a triage loop needs before it decides what to do.

For real wiring, run it as a server — the README's own pattern for putting decisions inside agentic workflows:
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v21 --port 8000
curl -s localhost:8000/v1/systemone \
-H 'content-type: application/json' \
-d '{
"state": "Help! My payouts have been failing for 3 days!",
"questions": {
"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
}
}'
The response carries the model name, the typed answers, token usage (note: one output token — nothing was generated), and the latency in milliseconds. From here the README's suggested uses read like an agentic-systems checklist: model routing (pick the right LLM for a task), tool selection (which tool the agent should call next), argument checking (verify a tool call before it runs), triage, guardrails, evaluations, and hybrid agents — the big LLM makes the hard decisions, the decider makes the rote ones. The repo ships a worked examples/strands/ example that gates a tool call with a before_tool_call intervention.

A ticket-triage decision pipeline with five moving parts: an install, a choice question that routes a ticket to billing with 0.89 of the distribution mass, a yes/no urgency check, a frustration score, all three batched into one call, and an HTTP endpoint your agent can query at ~115 ms a question. No JSON parsing, no prompt-format fragility, and a confidence number you can turn into an escalation policy.
--vision) needs the [vision] extra and transformers 5.18 or later; I left it out of this walkthrough deliberately.Still: for the rote decisions inside every agent — routing, tool selection, guardrails — a 2B model that answers in 115 ms with a confidence you can budget against is the cheapest upgrade on the board.