Stop paying flagship prices for easy questions: build an LLM cascade router that escalates only when it matters
Your API bill is dominated by easy questions answered by your most expensive model. Build a cascade router that tries the cheap model first and escalates only when confidence is low — pricing, backends, scorers, and thresholds, measured end-to-end.

Your most expensive model is answering your easiest questions. That is the quiet leak in almost every production LLM bill: extraction, classification, short factual answers, reformatting — the easy majority of traffic — all routed to the flagship model because sorting each question by hand is impractical. A cascade router fixes it mechanically: try the cheap model first, score its answer, and escalate to the flagship only when confidence is low.
Yesterday, OpenAI halved the usage allowance on its $200/month Pro plan. Whether you pay per seat or per token, the signal from every provider is the same: inference budgets are finite and the meter never stops. The academic answer to this problem is the LLM cascade. FrugalGPT (Chen, Zaharia, Zou, Stanford, 2023; arXiv:2305.05176, extended in TMLR 2024) showed you can match the best single model at up to 98% lower API cost: query the cheap model first, score its answer, escalate only when the score misses the threshold. In their HEADLINES case study the learned cascade ran GPT-J first, escalated to J1-L below a 0.96 score, and to GPT-4 below 0.37 — on one-fifth of GPT-4's budget.
In this tutorial you will build a production-shaped cascade router in Python: a priced model ladder, real provider backends with verified SDK calls, three confidence scorers, the routing loop itself, and a measurement harness that finds your threshold on a seeded, reproducible workload. Every number below was produced by running the code — including the places where the naive setup loses money.
What you'll need#
- Python 3.10 or later — check with
python3 --version. - The anthropic and openai Python packages (
pip install anthropic openai). Every SDK call in this article was verified against the installed packages (anthropic 1.9.0, openai 3.20.0) — parameter names included. - API keys in
ANTHROPIC_API_KEY(andOPENAI_API_KEYif you run the OpenAI backend). The measurement harness runs on deterministic mocks, so you can reproduce every chart with zero API spend and swap in the real backends later. - About 30 minutes. No GPU, no training, no labeled data required to start — though step 7 is better with your own eval set.
1. The math behind the cascade#
A cascade is a bet with a precise shape. For a two-tier setup (cheap, then strong), the expected cost per query is:
E[cost] = cost(cheap) + P(escalate) × cost(strong)
Notice the sting: on every escalation you pay for the cheap answer and throw it away. The cascade only wins when two things hold — the escalation probability is small, and the cheap tier is much cheaper than the strong one. Quality follows the same split:
accuracy = P(accept) × acc(cheap | accepted) + P(escalate) × acc(strong)
The threshold τ is the single dial that trades cost against quality, and the scorer's entire job is to make the accepted set coincide with the questions the cheap model actually gets right. A perfect scorer with τ at the cheap model's competence boundary gives you cheap prices on easy questions and flagship quality on hard ones. A bad scorer gives you the worst of both: you pay the cheap call, then pay the strong call anyway.
2. Price your ladder#
Everything downstream depends on real prices, so start there. These are Anthropic's published API rates, verified on platform.claude.com on September 29, 2026 (USD per million tokens). One pattern worth memorizing: output costs exactly 5× input on every Claude model.
| Model | Input / 1M | Output / 1M |
|---|---|---|
| Claude Haiku 4.5 | $1.00 | $5.00 |
| Claude Sonnet 5.5 | $2.00 | $10.00 |
| Claude Opus 5.5 | $4.00 | $20.00 |
Keep prices in a config dict, not in code — providers reprice, and a stale constant silently corrupts every cost report. The cost function is the only place prices should live:
from dataclasses import dataclass, field
# USD per 1M tokens. Verified 2026-09-29 on platform.claude.com.
# Re-check before budgeting: these move.
PRICING = {
"claude-haiku-4-5": {"in": 1.0, "out": 5.0},
"claude-sonnet-5-5": {"in": 2.0, "out": 10.0},
"claude-opus-5-5": {"in": 4.0, "out": 20.0},
}
def usd_cost(model: str, in_tokens: int, out_tokens: int) -> float:
p = PRICING[model]
return (in_tokens * p["in"] + out_tokens * p["out"]) / 1_000_000
At a typical short-answer profile (400 input tokens, 150 output tokens), one query costs $0.00115 on Haiku 4.5 versus $0.0046 on Opus 5.5 — a 4× spread. That spread is the entire economic opportunity the cascade harvests.
3. Wire up the backends#
Every backend exposes the same interface — ask(prompt) returns the text plus the token counts the provider billed — so the router never knows which provider it is talking to. The two real backends below use only call signatures verified against the installed SDKs. Two details I confirmed by reading the packages rather than the docs: anthropic 1.9.0's messages.create takes model, max_tokens, messages (and system) — there is no temperature parameter in this version, so don't pass one — and usage arrives as response.usage.input_tokens / output_tokens. On the OpenAI side, per-token logprobs require logprobs=True, top_logprobs=5 and live at choice.logprobs.content[i].logprob.
@dataclass
class Answer:
text: str
model: str
in_tokens: int
out_tokens: int
logprobs: list[float] = field(default_factory=list)
class ClaudeBackend:
def __init__(self, model: str, max_tokens: int = 512):
from anthropic import Anthropic
self.client = Anthropic() # reads ANTHROPIC_API_KEY
self.model, self.max_tokens = model, max_tokens
def ask(self, prompt: str, logprobs: bool = False) -> Answer:
msg = self.client.messages.create(
model=self.model,
max_tokens=self.max_tokens,
messages=[{"role": "user", "content": prompt}],
)
text = "".join(b.text for b in msg.content if b.type == "text")
return Answer(text, self.model,
msg.usage.input_tokens, msg.usage.output_tokens)
class OpenAIBackend:
def __init__(self, model: str, max_tokens: int = 512):
from openai import OpenAI
self.client = OpenAI() # reads OPENAI_API_KEY
self.model, self.max_tokens = model, max_tokens
def ask(self, prompt: str, logprobs: bool = False) -> Answer:
r = self.client.chat.completions.create(
model=self.model,
messages=[{"role": "user", "content": prompt}],
max_tokens=self.max_tokens,
temperature=0.0,
**({"logprobs": True, "top_logprobs": 5} if logprobs else {}),
)
ch = r.choices[0]
lps = ([t.logprob for t in ch.logprobs.content]
if (logprobs and ch.logprobs and ch.logprobs.content) else [])
return Answer(ch.message.content or "", self.model,
r.usage.prompt_tokens, r.usage.completion_tokens, lps)
The logprobs flag matters for the next step. Requesting logprobs is the cheapest confidence signal available — a few extra bytes on a response you already paid for.
4. Score the cheap answer#
The scorer is the cascade's brain, and you have three options in ascending order of cost:
1. Mean token logprob (free). Take the geometric mean of per-token probabilities — exp(mean(logprobs)) — as the confidence score. Fluent, on-rails answers score high; hesitant, low-probability token sequences score low. The catch: it only works where the provider exposes logprobs. OpenAI's chat completions do; Anthropic's Messages API does not. So this scorer is OpenAI-only in practice:
import math
def mean_logprob_scorer(answer: Answer) -> float:
"""Confidence in [0,1] from mean per-token logprob."""
if not answer.logprobs:
raise ValueError("needs per-token logprobs: call the backend with logprobs=True")
return math.exp(sum(answer.logprobs) / len(answer.logprobs))
2. A cheap judge model (one small extra call). This is the FrugalGPT pattern: their scorer was a DistilBERT fine-tuned for regression on (query, answer) pairs, outputting a reliability score in [0,1]. You don't need to train one — prompt a small model with the question, the candidate answer, and “reply with a single number 0–1 for correctness”. It costs one extra cheap call per candidate, works on any provider, and is the most reliable off-the-shelf option.
3. Self-reported confidence (free, weakest). Ask the model to end its answer with confidence: 0.x. It costs nothing extra, but models are systematically overconfident — treat a self-reported 0.9 as roughly 0.7 until you calibrate it on your own data. Use it only as a tiebreak, never as the gate.

5. Build the cascade loop#
The router itself is short. Walk each tier cheapest-first, score the answer, accept at or above the tier's threshold, and always accept at the final tier (threshold 0.0). The decision log is not optional decoration — it is your observability: escalation rate, score distribution, and per-tier spend all come from it.
@dataclass
class CascadeRouter:
tiers: list # backends, cheapest first
thresholds: list[float] # accept threshold per tier; last is 0.0
scorer = mean_logprob_scorer
log: list = field(default_factory=list)
def route(self, prompt: str, **ask_kwargs) -> Answer:
for i, backend in enumerate(self.tiers):
ans = backend.ask(prompt, **ask_kwargs)
score = 1.0 if i == len(self.tiers) - 1 else self.scorer(ans)
self.log.append({"model": backend.model,
"score": round(score, 4),
"accepted": score >= self.thresholds[i],
"cost_usd": usd_cost(backend.model,
ans.in_tokens, ans.out_tokens)})
if score >= self.thresholds[i]:
return ans
return ans # unreachable: the last threshold is 0.0
def total_cost(self) -> float:
return sum(e["cost_usd"] for e in self.log)
Two design decisions are load-bearing. First, thresholds are per-tier, not global — the middle of the ladder can be stricter or looser than the bottom. Second, the final tier always accepts, so the router can never return nothing; a cascade must degrade to the flagship, never to an error.
6. Measure it on a workload#
A cascade tuned on vibes is a random number generator with extra steps. The harness below runs the router over 500 seeded questions with known difficulty: the mock cheap model is correct on the easy 75%, the mock strong model on 97%, and the logprob scorer genuinely separates right from wrong answers (scores sag as difficulty rises, with noise — like the real thing). Token profile is 400 in / 150 out per query. This is a simulated workload — your absolute numbers will differ — but the method is exactly what you run with real backends and your own labeled eval set.
python3 cascade_router.py
== headline run: 2-tier cascade, thresholds [0.42, 0.0] ==
{
"n": 500,
"accuracy": 0.966,
"cost_usd": 1.1914,
"baseline_usd": 2.3,
"savings_pct": 48.2,
"tier_share": {"claude-haiku-4-5": 0.732, "claude-opus-5-5": 0.268},
"strong_accuracy": 0.97
}
Read it as: the cheap model answered 73.2% of questions, the flagship took the other 26.8%, total accuracy 0.966 versus 0.97 for always-flagship — and the bill fell 48.2%, from $2.30 to $1.19 per 500 queries. At a million queries a month, that is the difference between $4,600 and $2,383.
The harness also exposes the classic cascade failure — the dead middle tier. A three-tier run (cheap → mid → strong, thresholds [0.42, 0.50, 0.0]) saved only 34.8%, and the mid tier accepted 0% of questions. Every tier you add taxes every escalation with another wasted call, so a tier that never accepts is pure overhead. Check tier_share in your logs: if a tier rounds to zero, delete it.
7. Tune the threshold#
The threshold is not a hyperparameter you guess — you sweep it and read the curve. Same 500 questions, varying only the cheap tier's τ:
tau=0.20 accuracy=0.760 savings=75.0% cheap answers everything (including wrong)
tau=0.30 accuracy=0.770 savings=72.8%
tau=0.42 accuracy=0.966 savings=48.2% <- sweet spot
tau=0.55 accuracy=0.966 savings=25.6%
tau=0.70 accuracy=0.966 savings=6.4%
tau=0.85 accuracy=0.966 savings=-10.6% threshold so strict you pay for the
cheap call, then escalate anyway

Two cliffs, one plateau. Below τ≈0.4 the gate accepts answers the cheap model got wrong and accuracy falls off a cliff. Above τ≈0.8 the gate rejects good cheap answers and savings go negative — you paid for the cheap call and escalated anyway, the worst possible outcome. The plateau between them is where you live. In production, the procedure is: fix your minimum acceptable accuracy, take the cheapest τ that holds it on a held-out eval set, and re-tune monthly — traffic mixes drift, and a threshold tuned on last quarter's questions is a guess about this quarter's.
8. Harden it for production#
The loop above is the idea; production needs five more things:
- Timeouts with escalation-as-fallback. If the cheap call times out or errors, don't fail the request — escalate. If every tier fails, fail loudly. A cascade must never silently serve nothing.
- An exact-match cache on (prompt, tier) in front of the router. Repeated questions should never reach any model twice.
- Concurrency. The loop is per-request and sequential by design; run requests in parallel with
asyncioand a semaphore sized to your provider's rate limits. - Drift alerts on the decision log. Track escalation rate and score distribution per tier. A sudden jump means your traffic mix changed or the cheap model got worse — either way, your threshold is now wrong.
- A per-request budget cap. If cumulative spend would exceed the cap, stop escalating and return the best answer so far. This bounds the worst case no matter what the scorer does.
Which approach should you use?#
Cascade (this tutorial) wins when question difficulty varies and you can score answers after the fact. Its tax is the wasted cheap call on every escalation.
Pre-routing (the RouteLLM pattern) predicts difficulty from the prompt before calling anything — no wasted calls. LMSYS's RouteLLM (ICLR 2025, “Learning to Route LLMs with Preference Data”) learns P(strong model beats weak model | query) from preference data with routers ranging from similarity-weighted ranking to fine-tuned classifiers. Choose it when you have win-rate training data and want to skip the cheap call entirely; choose the cascade when you don't.
Prompt caching is orthogonal, not alternative: repeated prefixes (system prompts, documents) are served from cache at 10% of input price on Anthropic. Stack it under the cascade — cache the cheap tier's prompt and the flagship's prompt independently.
Batch API gives 50% off both input and output for async workloads on Anthropic. If your workload tolerates hours of latency, batch first and cascade second; the discounts multiply.
When not to cascade: uniformly hard traffic (everything escalates — you bought latency for nothing), no reliable way to score answers, or latency SLOs tighter than two sequential model calls. Measure your difficulty mix first; the cascade is an economic structure, not a default.
The takeaway#
A cascade router is not a prompt trick — it is an economic structure wrapped around your models. Price the ladder from the provider's page, not from memory. Pick the cheapest scorer your provider supports (logprobs where they exist, a small judge everywhere else). Sweep the threshold on a labeled set instead of guessing it, and keep the decision log — escalation rate is the vital sign of the whole system. On a mixed workload, the measured result here was 48% cheaper at flagship accuracy, with the cheap model quietly handling nearly three-quarters of traffic. The flagship is still there for the questions that deserve it. It just stopped answering the ones that don't.