
Every few months the same argument erupts online: the model got worse. Quieter, cheaper, dumber — "nerfed" a few weeks after launch, right after everyone locked in their subscriptions. The counterargument is always the same too: you're pattern-matching on noise, nothing changed. It's vibes versus vibes, because nobody has a day-zero baseline to check against.
livenerf (ninjahawk/livenerf, 1,200+ stars, MIT) decided to stop arguing and start measuring. It's a small, append-only benchmark for exactly one question — does a frontier model get worse after it ships? — and it started the clock on launch day: September 22, 2026, the day Claude Opus 5.5 came out. One frozen panel of questions, one run a day, thirty days, with the statistics pre-registered before the first result.
You don't need livenerf's specific setup to use its playbook. The method transfers to any model you depend on: an API model whose provider ships silent updates, a local model you re-quantize, a fine-tune you iterate on. Freeze the questions, pin everything else, run daily, and decide with paired statistics instead of your eyes. Here's the whole thing, hands-on.
pip install inspect_ai pandas — Inspect is the UK AI Safety Institute's open-source eval framework, and it's what livenerf itself is built on (it pins inspect_ai==0.3.266).This is the core insight, and it's the opposite of what you'd guess. livenerf screened 2,336 questions — GPQA Diamond, MMLU-Pro, competition math, AIME 2025–26 — running each one four times. Opus 5.5 got about 93% right on the first try, and 97% of the questions were always right or always wrong. It kept the 78 "sometimes right" ones and threw the rest away.
Why: questions the model always gets right can't detect a decline, and questions it always gets wrong can't either. Only the borderline questions move when capability moves. A panel of 50 questions your model aces is a smoke detector with the battery removed.
So: assemble your candidates, run each a few times, and keep the ones with mixed results. Then freeze the list — livenerf stores its panel in data/standard_panel.json with a lock file, and the daily run refuses to start if the panel changed. Write the task in Inspect's task API:
# drift/tasks.py
import json
from inspect_ai import Task, task
from inspect_ai.dataset import MemoryDataset, Sample
from inspect_ai.model import GenerateConfig
from inspect_ai.solver import generate, system_message
SYSTEM_PROMPT = "Answer each question. Put your final answer inside <answer>...</answer>."
PANEL_VERSION = "v1" # bump this and your old baseline is void
@task
def drift_panel() -> Task:
items = json.load(open("drift/panel.json")) # [{"id","input","target"}]
dataset = MemoryDataset(
[Sample(id=it["id"], input=it["input"], target=it["target"]) for it in items],
name="drift-panel",
)
return Task(
dataset=dataset,
solver=[system_message(SYSTEM_PROMPT), generate()],
scorer=choice_scorer(),
name="drift_panel",
config=GenerateConfig(),
version=f"panel-{PANEL_VERSION}/sys-v1",
fail_on_error=False,
)
The grader must be a pure function — no I/O, no model calls, no randomness, same input and same score forever. livenerf's multiple-choice grader is eleven lines and worth copying verbatim:
def choice_score(answer, target):
"""1.0 if the answer is exactly the target letter
(case-insensitive, optional parentheses)."""
if answer is None:
return 0.0
return float(answer.strip().strip("()").strip().upper() == target)
Wrap it as an Inspect scorer (this mirrors livenerf/scorers/__init__.py, which also attaches mean() and item-clustered stderr() metrics for the error bars later):
from inspect_ai.scorer import Score, Target, mean, scorer, stderr
from inspect_ai.solver import TaskState
import re
ANSWER_RE = re.compile(r"<answer>(.*?)</answer>", re.DOTALL | re.IGNORECASE)
@scorer(metrics=[mean(), stderr()])
def choice_scorer():
async def score(state: TaskState, target: Target) -> Score:
found = ANSWER_RE.findall(state.output.completion or "")
answer = found[-1].strip() if found else None
return Score(value=choice_score(answer, target.text), answer=answer)
return score

A drift monitor measures one variable: the model. Everything else must be frozen, or a score change means nothing. livenerf pins the system prompt version, the harness (it records the git hash of the harness code with every run), the grader, and the CLI version — its daily entry point checks the installed CLI against a pinned version file and refuses to run on a mismatch.
Do the same in your run metadata. Record the model id, the panel version, the Inspect version, and the system prompt version with every run, and make your runner abort if the panel file changed since the baseline started. The discipline is the product: a score you can't attribute is just a number.
Run the panel once to establish day zero. From the CLI, the pattern is the same one livenerf documents for its own tasks:
inspect eval drift/tasks.py@drift_panel --log-dir logs
Or from Python, following livenerf/daily.py's entry point (note the provider/model id format and the metadata that makes each run attributable):
from inspect_ai import eval as inspect_eval
from drift.tasks import drift_panel
logs = inspect_eval(
[drift_panel()],
model="YOUR_PROVIDER/YOUR_MODEL", # e.g. the model you ship against
log_dir="logs",
tags=["drift-watch", "daily"],
metadata={"panel_version": "v1", "baseline_day": "2026-10-04"},
display="none",
max_connections=4,
)
Inspect writes .eval logs into logs/. Extract each day's per-question scores and append one line per day to logs/daily.jsonl — date, score per question, total output tokens per question. Tokens matter; step 6 explains why.
livenerf's daily run fires from cron every hour but does nothing once today's run is recorded, so a blocked or missed hour catches up at the next one instead of being lost. Steal that shape:
# crontab: try hourly during the day; the script no-ops if today already ran
7 6-22 * * * cd /path/to/drift-watch && bash scripts/daily.sh >> logs/daily.log 2>&1
#!/usr/bin/env bash
# scripts/daily.sh
set -euo pipefail
cd "$(dirname "$0")/.."
python -m drift.daily # checks panel lock + model pin, runs once, appends to logs/daily.jsonl
livenerf also guards on cost: its runner skips (and retries an hour later) when its usage meters are at or above caps (--weekly-cap 75, --five-hour-cap 60 by default). Your version of that is a budget check before the run — a monitor that bankrupts you is worse than vibes.
One day's score is noise. livenerf's design: days 1–10 form the baseline, then two 10-day decision windows, with the first possible call around October 24. The decision uses paired day-vs-baseline deltas with item-clustered standard errors, following Anthropic's "Adding Error Bars to Evals" — nothing homebrew to argue about. It can detect an accuracy change of about 7.5 points per 10-day window.
Your minimal version: after ten baseline days, compare each new day's item-level scores against the baseline with a paired test. In pandas, roughly:
import json, pandas as pd
days = [json.loads(l) for l in open("logs/daily.jsonl")]
scores = pd.DataFrame([d["scores"] for d in days], index=[d["date"] for d in days])
base = scores.iloc[:10] # the frozen baseline window — never move it
daily_mean = scores.mean(axis=1)
base_mean = base.mean().mean()
base_se = base.mean(axis=1).std() / (len(base) ** 0.5)
today = daily_mean.iloc[-1]
print(f"baseline {base_mean:.1%} ± {1.96*base_se:.1%} | today {today:.1%}")
print("DRIFT?", abs(today - base_mean) > 1.96 * base_se)
That's a sketch, not the method — the rigorous version clusters standard errors by question and works on paired deltas, which is what livenerf's livenerf.analysis module does. But the shape is right: a fixed baseline window, a pre-chosen decision rule, and no moving the goalposts after you see the numbers. If you clone livenerf itself, python -m livenerf.plot redraws the daily score chart (light and dark SVG, no dependencies) straight from the .eval logs.

livenerf's validation runs produced the single most useful finding for anyone building a monitor: when the model's effort drops, it shows up in token counts long before it shows up in accuracy. Forcing low effort cut output tokens by 62% while accuracy fell only 8.3 ± 4.5 points; medium effort cut tokens 26% for 4.2 ± 3.9 points of accuracy. A provider quietly routing you to a cheaper configuration looks, at first, like your model got terser — not dumber.
So log output tokens per question alongside scores, and alert on token drift too. It's the canary.
A drift watchdog with the livenerf shape: a frozen panel of borderline questions, a pinned harness, one run a day appended to a JSONL log, and a pre-registered statistical rule that turns "the model feels worse" into a yes-or-no answer. As of October 3, the real livenerf is 10 of 30 days in with zero missed days — baseline complete, first decision window open, first possible call around October 24. Your version starts its own clock today.