Stop vibe-checking your model: write real evals with inspect_ai, the UK AI Safety Institute's framework
You changed the system prompt, tried three questions by hand, squinted at the answers, and shipped it. That is how most teams "evaluate" their LLM features — and it is why most of them cannot tell you whether last week's prompt edit helped or hurt. The fix is a real eval harness: fixed datasets, scripted graders, and logs you can diff instead of vibes you forget. The UK AI Safety Institute built that harness, open source, and it has quietly become the default tool serious eval work is written in: inspect_ai. In this tutorial you build three working evals in it — a multiple-choice benchmark, a custom numeric grader, and a full tool-using agent eval — and every line runs on your machine with zero API keys, because inspect_ai ships a scriptable mock model built for exactly this.

Here is the uncomfortable truth about LLM development in 2026: the models got dramatically better, and most teams' measurement discipline did not. A prompt tweak here, a model swap there, a thumbs-up from a product manager — and nobody can reconstruct, a month later, what was actually tried or what it did. Frontier labs don't work this way. They work in evals: versioned datasets of test cases, deterministic graders, and logs that make every run comparable to every other run.
inspect_ai is the framework the UK AI Safety Institute (built with Meridian Labs) wrote to do that work at scale. It is one mental model with three moving parts:
- Dataset — test cases, each carrying an
inputand atarget. - Solver — whatever produces an answer: one model call, a prompt chain, or a full tool-using agent.
- Scorer — whatever turns an answer into a number: text matching, a custom function, or another model grading.
A Task binds the three together. Around that core you get the infrastructure that makes evals livable: 20+ model providers, sandboxed tool execution, parallel runs, resumable logs, and a web viewer for reading transcripts. There are also 200+ prebuilt benchmark implementations in the companion inspect_evals package — but this tutorial is about writing your own, because the eval that matters to you is the one that measures your product.
What you'll need
- Python 3.10+ and
pip install inspect-ai. The reference run used inspect_ai 0.3.273 on Python 3.12 — CPU only, no GPU anywhere in this tutorial. - Zero API keys. Every eval below runs against
mockllm/model, inspect_ai's built-in mock provider. You can script exactly what it returns, so the scores below are deterministic and reproducible on your machine. When you have a real key, swapping inopenai/gpt-4o-minioranthropic/claude-sonnet-4-6is a one-line change — I show where. - About 20 minutes. Three small files, each one runnable top to bottom. Every number quoted below came from an actual run, and I'll show you the exact outputs.
1. Your first eval in twenty lines
The smallest useful eval: four multiple-choice questions, a solver that formats them, a scorer that checks the letter. Save this as hello_eval.py:
from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import choice
from inspect_ai.solver import multiple_choice
QUESTIONS = [
("What is the capital of France?",
["London", "Berlin", "Paris", "Madrid"], "C"),
("2 + 2 * 2 = ?", ["4", "6", "8", "2"], "B"),
("Which planet is known as the Red Planet?",
["Venus", "Mars", "Jupiter", "Saturn"], "B"),
("What does CPU stand for?",
["Central Process Unit", "Central Processing Unit",
"Computer Personal Unit", "Central Processor Utility"], "B"),
]
@task
def capitals():
return Task(
dataset=[
Sample(input=q, choices=choices, target=target)
for q, choices, target in QUESTIONS
],
solver=[multiple_choice()],
scorer=choice(),
)
if __name__ == "__main__":
# Script the mock model: it "answers" C, B, A, B.
# 3 of 4 correct on purpose, so we get a real 0.75, not a toy 1.0.
mock_answers = ["ANSWER: C", "ANSWER: B", "ANSWER: A", "ANSWER: B"]
log = eval(
capitals(),
model="mockllm/model",
model_args={
"custom_outputs": (
ModelOutput.from_content(model="mockllm", content=a)
for a in mock_answers
)
},
)[0]
print("accuracy:", log.results.scores[0].metrics["accuracy"].value)
Run it with python hello_eval.py. Output:
choice
accuracy 0.750
stderr 0.250
Log: logs/2026-10-01T23-31-06-00-00_capitals_jaRoNHsqQ8Ng322hHF9okj.eval
accuracy: 0.75
Three things are worth noticing, because they are the whole framework in miniature:
- The solver owns the response format.
multiple_choice()doesn't just paste the question — it instructs the model to answer asANSWER: $LETTER, andchoice()parses exactly that. When I first scripted the mock to return bare letters likeC, every sample scoredinvalid_response_format. Real models fail this way too, and now you know where that failure lives: in the solver/scorer contract, not in the model. mockllm/modelis a first-class testing tool. Itscustom_outputsmodel arg accepts an iterable (or generator, or callable) ofModelOutputobjects. That one hook is what makes this entire tutorial keyless: you script the model's behavior, and the eval machinery — formatting, scoring, logging — runs for real.- The stderr is the point. 0.75 ± 0.25 on four samples is a joke of a measurement, and the framework says so to your face. A serious eval has hundreds of samples; the machinery is identical.
With a real key, the same file runs against a frontier model two ways — CLI (inspect eval hello_eval.py --model openai/gpt-4o-mini) or the Python eval() call you already used, with model="openai/gpt-4o-mini".

2. Read the log like a researcher
Every run writes an .eval log into ./logs/. This is the artifact the whole framework is organized around: per-sample inputs, the full message transcript, token usage, and scores — in one versioned, diffable file. Two ways to read it:
inspect view starts a local web UI for browsing runs, drilling into transcripts, and comparing samples. Programmatically, the log is just data:
from inspect_ai.log import read_eval_log
log = read_eval_log("logs/2026-10-01T23-31-06-00-00_capitals_jaRoNHsqQ8Ng322hHF9okj.eval")
s = log.samples[2] # the one we got wrong
print("TARGET:", s.target) # B
print("COMPLETION:", s.output.completion) # ANSWER: A
print("SCORE:", s.scores["choice"].value) # I (incorrect)
This is the habit that separates evals from vibes: when the scoreboard says 0.75, you don't re-run and hope — you open the log and read the one failure. Samples carry metadata too, so you can slice scores by category (question type, difficulty, model version) with grouped metrics later. Treat the log as the deliverable and the terminal scoreboard as the summary.
3. Write your own scorer
Built-in scorers (choice, includes, match, model_graded_qa) cover common shapes. Everything else is a @scorer function that takes the task state and the target and returns a Score. Save this as math_eval.py:
from inspect_ai import Task, eval, task
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import Score, Target, mean, scorer, stderr
from inspect_ai.solver import generate
from inspect_ai.solver._task_state import TaskState
@scorer(metrics=[mean(), stderr()])
def numeric_within_tolerance(tolerance: float = 0.05):
async def score(state: TaskState, target: Target):
raw = state.output.completion.strip().replace(",", "")
try:
answer = float(raw)
expected = float(target.text)
correct = abs(answer - expected) <= tolerance
except (ValueError, TypeError):
answer, correct = None, False
return Score(
value=1.0 if correct else 0.0,
answer=str(answer),
explanation=(
f"model answered {answer}, expected {target.text} "
f"(tolerance {tolerance})"
),
)
return score
@task
def mental_math():
return Task(
dataset=[
Sample(input="What is 15% of 240? Reply with just the number.",
target="36"),
Sample(input="What is 7 * 13? Reply with just the number.",
target="91"),
Sample(input="What is 144 / 12? Reply with just the number.",
target="12"),
Sample(input="What is 2^10? Reply with just the number.",
target="1024"),
],
solver=[generate()],
scorer=numeric_within_tolerance(),
)
if __name__ == "__main__":
# 2 right, 1 wrong, 1 non-numeric (exercises the parse-failure path)
mock_answers = ["36", "90", "twelve", "1024"]
log = eval(
mental_math(),
model="mockllm/model",
model_args={
"custom_outputs": (
ModelOutput.from_content(model="mockllm", content=a)
for a in mock_answers
)
},
)[0]
print("mean score:", log.results.scores[0].metrics["mean"].value)
for s in log.samples:
sc = s.scores["numeric_within_tolerance"]
print(f" score={sc.value} explanation={sc.explanation}")
Output:
numeric_within_tolerance
mean 0.500
stderr 0.289
mean score: 0.5
score=1.0 explanation=model answered 36.0, expected 36 (tolerance 0.05)
score=0.0 explanation=model answered 90.0, expected 91 (tolerance 0.05)
score=0.0 explanation=model answered None, expected 12 (tolerance 0.05)
score=1.0 explanation=model answered 1024.0, expected 1024 (tolerance 0.05)
The anatomy of a scorer, all in one function:
state.output.completionis the model's final text;target.textis the expected answer. The scorer is async and can do anything — call a unit test, run a regex, or invoke another model (that's whatmodel_graded_qais).Score(value=..., answer=..., explanation=...)separates the number (what aggregates) from the human-readable record (what you read in the log). Always fill inexplanation— future-you debugging a regression will thank present-you.metrics=[mean(), stderr()]declares how per-sample scores aggregate. Custom metrics (pass@k, grouped-by-metadata) plug in here.- Graders must decide what garbage means. The model answered
twelve; the scorer fails it closed withanswer=None. That decision is now explicit, versioned code — not a judgment call you make differently each Tuesday.
4. Evaluate a tool-using agent
This is inspect_ai's sweet spot — the thing a 30-line script can't do. You give the model a real tool, run the built-in react() agent (reason → act → observe → submit), and score the transcript: did the agent actually call the tool with the right arguments?
First the tool. Custom tools are factory functions — the framework builds the tool's schema from the inner function's signature and docstring, and every parameter needs a documented description (I hit that validation error so you don't have to):
from inspect_ai.tool import Tool, tool
@tool
def multiply() -> Tool:
async def execute(a: float, b: float):
"""Multiply two numbers together. Always use this tool for
multiplication instead of doing the arithmetic yourself.
Args:
a: the first number to multiply
b: the second number to multiply
"""
return a * b
return execute
Now the agent eval. The scorer inspects tool_calls on the assistant messages — the structured record of what the agent did, not the text it wrote:
from inspect_ai import Task, eval, task
from inspect_ai.agent import react
from inspect_ai.dataset import Sample
from inspect_ai.model import ModelOutput
from inspect_ai.scorer import Score, Target, mean, scorer, stderr
from inspect_ai.solver._task_state import TaskState
from inspect_ai.tool import Tool, tool
# The multiply tool from the previous step, included so this file
# runs on its own:
@tool
def multiply() -> Tool:
async def execute(a: float, b: float):
"""Multiply two numbers together. Always use this tool for
multiplication instead of doing the arithmetic yourself.
Args:
a: the first number to multiply
b: the second number to multiply
"""
return a * b
return execute
@scorer(metrics=[mean(), stderr()])
def used_tool_correctly():
async def score(state: TaskState, target: Target):
calls = []
for m in state.messages:
for tc in getattr(m, "tool_calls", None) or []:
if tc.function == "multiply":
calls.append(tc)
hit = any(
tc.arguments.get("a") == 7 and tc.arguments.get("b") == 6
for tc in calls
)
return Score(
value=1.0 if hit else 0.0,
answer=f"{len(calls)} multiply call(s)",
explanation=f"multiply(7, 6) called with correct args: {hit}",
)
return score
@task
def calculator_agent():
return Task(
dataset=[
Sample(
input="What is 7 times 6? Use the multiply tool, "
"then report the answer.",
target="42",
),
],
solver=[react(
prompt="You are a calculator agent.",
tools=[multiply()],
attempts=2,
)],
scorer=used_tool_correctly(),
)
if __name__ == "__main__":
log = eval(
calculator_agent(),
model="mockllm/model",
model_args={
"custom_outputs": [
# turn 1: the "model" calls the tool with correct args
ModelOutput.for_tool_call(
model="mockllm", tool_name="multiply",
tool_arguments={"a": 7, "b": 6},
),
# turn 2: it sees the result, then answers in text
ModelOutput.from_content(
model="mockllm", content="The answer is 42."),
# turn 3: react() ends the loop on a submit() call
ModelOutput.for_tool_call(
model="mockllm", tool_name="submit",
tool_arguments={"answer": "42"},
),
]
},
)[0]
print("mean score:", log.results.scores[0].metrics["mean"].value)
Output: mean score: 1.0. And the transcript in the log is the real prize — the complete reason-act-observe loop, executed by actual code:
== user
What is 7 times 6? Use the multiply tool, then report the answer.
== assistant
tool call for tool multiply
== tool
42.0
== assistant
The answer is 42.
== user
Please proceed to the next step... call the `submit()` tool with your final answer.
== assistant
tool call for tool submit
42
The mock scripted the model's side, but everything else ran for real: the tool schema was generated from your docstring, the framework dispatched the call, your Python execute ran and returned 42.0, and the scorer verified the call arguments from the transcript. Swap mockllm/model for a real model and the only thing that changes is who decides to call the tool.
Two details that matter in production:
- The loop ends at
submit().react()appends asubmit(answer: str)tool; the agent runs until the model calls it. Withattempts=2, the agent even re-scores its own submission with your scorer and retries when it isn't 1.0 — your grader is load-bearing inside the loop, not just after it. - Score behavior, not text. A model that writes "42" without touching the tool fails this eval — correctly. For agent evals, the transcript is the answer.

Which approach should you use?
inspect_ai is the right default for serious eval work, but it isn't the only tool. Pick by the shape of your problem:
| Your situation | Use | Why |
|---|---|---|
| A 30-line one-off check you'll run twice | A plain script | inspect_ai's structure is overhead for trivial evals — its own docs say so. |
| Custom agent or tool-use evals you need reproducible and shareable | inspect_ai | Solvers, sandboxed tools, transcripts-as-evidence, and logs you can diff. This tutorial. |
| Academic log-probability benchmarks (MMLU, ARC, HellaSwag) with established numbers | lm-evaluation-harness (EleutherAI) | The comparable-scores tooling lives there; inspect_ai is generation-based. |
| Prompt regression tests running in CI on every commit | promptfoo | Purpose-built for prompt-as-code workflows — see our hands-on promptfoo tutorial. |
| "Just have an LLM grade the outputs" | Fix the process first | Model-graded scoring is a component (model_graded_qa), not an eval strategy. Read our guide to red-teaming your LLM judges before you trust one. |
Where to go next inside inspect_ai: the inspect_evals package ships 200+ ready-to-run benchmarks (SWE-bench, GPQA, CyBench, AgentHarm among them) — run one against your model before writing your own, to calibrate. Then write the eval only you can write: the one whose samples are your product's actual failure modes. An eval of twenty well-chosen real failures beats a benchmark of two thousand generic questions.
The takeaway
Evals are the difference between "we think the new prompt is better" and "the new prompt scores 0.81 ± 0.03 against 0.74 ± 0.03 on the same 400 cases, and here is the log." inspect_ai gives you the three primitives — dataset, solver, scorer — plus the infrastructure that makes them livable: scriptable mock models for keyless development, real tool execution inside agent loops, transcript-aware graders, and logs that are the deliverable. You just ran all of it with no API key and no GPU. The next step is the one that actually matters: point it at your own product's failure modes, and never ship on vibes again.
Every number in this article was measured, not sketched: inspect_ai 0.3.273 on Python 3.12, CPU only, scripted mockllm outputs — multiple-choice accuracy 0.75 ± 0.25 (n=4), custom numeric scorer mean 0.50 ± 0.29 (n=4), agent tool-use eval 1.0 (n=1). The three scripts run top to bottom as printed.