AI Frontier Post
A declining benchmark score line over 30 days with a confidence band and a dashed day-0 baseline
The shape of a drift experiment: a daily score, a frozen baseline, and the band that decides whether the model moved. Illustration: AI Frontier Post.

Every few months the same argument erupts online: the model got worse. Quieter, cheaper, dumber — "nerfed" a few weeks after launch, right after everyone locked in their subscriptions. The counterargument is always the same too: you're pattern-matching on noise, nothing changed. It's vibes versus vibes, because nobody has a day-zero baseline to check against.

livenerf (ninjahawk/livenerf, 1,200+ stars, MIT) decided to stop arguing and start measuring. It's a small, append-only benchmark for exactly one question — does a frontier model get worse after it ships? — and it started the clock on launch day: September 22, 2026, the day Claude Opus 5.5 came out. One frozen panel of questions, one run a day, thirty days, with the statistics pre-registered before the first result.

You don't need livenerf's specific setup to use its playbook. The method transfers to any model you depend on: an API model whose provider ships silent updates, a local model you re-quantize, a fine-tune you iterate on. Freeze the questions, pin everything else, run daily, and decide with paired statistics instead of your eyes. Here's the whole thing, hands-on.

What you'll need

1. Freeze a panel of questions the model gets sometimes right

This is the core insight, and it's the opposite of what you'd guess. livenerf screened 2,336 questions — GPQA Diamond, MMLU-Pro, competition math, AIME 2025–26 — running each one four times. Opus 5.5 got about 93% right on the first try, and 97% of the questions were always right or always wrong. It kept the 78 "sometimes right" ones and threw the rest away.

Why: questions the model always gets right can't detect a decline, and questions it always gets wrong can't either. Only the borderline questions move when capability moves. A panel of 50 questions your model aces is a smoke detector with the battery removed.

So: assemble your candidates, run each a few times, and keep the ones with mixed results. Then freeze the list — livenerf stores its panel in data/standard_panel.json with a lock file, and the daily run refuses to start if the panel changed. Write the task in Inspect's task API:

# drift/tasks.py
import json
from inspect_ai import Task, task
from inspect_ai.dataset import MemoryDataset, Sample
from inspect_ai.model import GenerateConfig
from inspect_ai.solver import generate, system_message

SYSTEM_PROMPT = "Answer each question. Put your final answer inside <answer>...</answer>."
PANEL_VERSION = "v1"  # bump this and your old baseline is void

@task
def drift_panel() -> Task:
    items = json.load(open("drift/panel.json"))  # [{"id","input","target"}]
    dataset = MemoryDataset(
        [Sample(id=it["id"], input=it["input"], target=it["target"]) for it in items],
        name="drift-panel",
    )
    return Task(
        dataset=dataset,
        solver=[system_message(SYSTEM_PROMPT), generate()],
        scorer=choice_scorer(),
        name="drift_panel",
        config=GenerateConfig(),
        version=f"panel-{PANEL_VERSION}/sys-v1",
        fail_on_error=False,
    )

The grader must be a pure function — no I/O, no model calls, no randomness, same input and same score forever. livenerf's multiple-choice grader is eleven lines and worth copying verbatim:

def choice_score(answer, target):
    """1.0 if the answer is exactly the target letter
    (case-insensitive, optional parentheses)."""
    if answer is None:
        return 0.0
    return float(answer.strip().strip("()").strip().upper() == target)

Wrap it as an Inspect scorer (this mirrors livenerf/scorers/__init__.py, which also attaches mean() and item-clustered stderr() metrics for the error bars later):

from inspect_ai.scorer import Score, Target, mean, scorer, stderr
from inspect_ai.solver import TaskState
import re

ANSWER_RE = re.compile(r"<answer>(.*?)</answer>", re.DOTALL | re.IGNORECASE)

@scorer(metrics=[mean(), stderr()])
def choice_scorer():
    async def score(state: TaskState, target: Target) -> Score:
        found = ANSWER_RE.findall(state.output.completion or "")
        answer = found[-1].strip() if found else None
        return Score(value=choice_score(answer, target.text), answer=answer)
    return score
A terminal running a nightly LLM evaluation with per-question scores
Illustration: the nightly run — every question in the frozen panel, asked once a day, graded by a pure function. AI-generated.

2. Pin everything except the questions

A drift monitor measures one variable: the model. Everything else must be frozen, or a score change means nothing. livenerf pins the system prompt version, the harness (it records the git hash of the harness code with every run), the grader, and the CLI version — its daily entry point checks the installed CLI against a pinned version file and refuses to run on a mismatch.

Do the same in your run metadata. Record the model id, the panel version, the Inspect version, and the system prompt version with every run, and make your runner abort if the panel file changed since the baseline started. The discipline is the product: a score you can't attribute is just a number.

3. Run your day-zero baseline

Run the panel once to establish day zero. From the CLI, the pattern is the same one livenerf documents for its own tasks:

inspect eval drift/tasks.py@drift_panel --log-dir logs

Or from Python, following livenerf/daily.py's entry point (note the provider/model id format and the metadata that makes each run attributable):

from inspect_ai import eval as inspect_eval
from drift.tasks import drift_panel

logs = inspect_eval(
    [drift_panel()],
    model="YOUR_PROVIDER/YOUR_MODEL",   # e.g. the model you ship against
    log_dir="logs",
    tags=["drift-watch", "daily"],
    metadata={"panel_version": "v1", "baseline_day": "2026-10-04"},
    display="none",
    max_connections=4,
)

Inspect writes .eval logs into logs/. Extract each day's per-question scores and append one line per day to logs/daily.jsonl — date, score per question, total output tokens per question. Tokens matter; step 6 explains why.

4. Schedule it, with a no-op guard

livenerf's daily run fires from cron every hour but does nothing once today's run is recorded, so a blocked or missed hour catches up at the next one instead of being lost. Steal that shape:

# crontab: try hourly during the day; the script no-ops if today already ran
7 6-22 * * * cd /path/to/drift-watch && bash scripts/daily.sh >> logs/daily.log 2>&1
#!/usr/bin/env bash
# scripts/daily.sh
set -euo pipefail
cd "$(dirname "$0")/.."
python -m drift.daily   # checks panel lock + model pin, runs once, appends to logs/daily.jsonl

livenerf also guards on cost: its runner skips (and retries an hour later) when its usage meters are at or above caps (--weekly-cap 75, --five-hour-cap 60 by default). Your version of that is a budget check before the run — a monitor that bankrupts you is worse than vibes.

5. Read the drift with statistics, not your eyes

One day's score is noise. livenerf's design: days 1–10 form the baseline, then two 10-day decision windows, with the first possible call around October 24. The decision uses paired day-vs-baseline deltas with item-clustered standard errors, following Anthropic's "Adding Error Bars to Evals" — nothing homebrew to argue about. It can detect an accuracy change of about 7.5 points per 10-day window.

Your minimal version: after ten baseline days, compare each new day's item-level scores against the baseline with a paired test. In pandas, roughly:

import json, pandas as pd

days = [json.loads(l) for l in open("logs/daily.jsonl")]
scores = pd.DataFrame([d["scores"] for d in days], index=[d["date"] for d in days])
base = scores.iloc[:10]            # the frozen baseline window — never move it

daily_mean = scores.mean(axis=1)
base_mean = base.mean().mean()
base_se = base.mean(axis=1).std() / (len(base) ** 0.5)
today = daily_mean.iloc[-1]
print(f"baseline {base_mean:.1%} ± {1.96*base_se:.1%} | today {today:.1%}")
print("DRIFT?", abs(today - base_mean) > 1.96 * base_se)

That's a sketch, not the method — the rigorous version clusters standard errors by question and works on paired deltas, which is what livenerf's livenerf.analysis module does. But the shape is right: a fixed baseline window, a pre-chosen decision rule, and no moving the goalposts after you see the numbers. If you clone livenerf itself, python -m livenerf.plot redraws the daily score chart (light and dark SVG, no dependencies) straight from the .eval logs.

A 30-day calendar where each day's quiz card is graded, forming a declining trend line
Illustration: thirty days, one frozen panel, one score a day — the shape of a real drift experiment. AI-generated.

6. Watch the tokens — they move first

livenerf's validation runs produced the single most useful finding for anyone building a monitor: when the model's effort drops, it shows up in token counts long before it shows up in accuracy. Forcing low effort cut output tokens by 62% while accuracy fell only 8.3 ± 4.5 points; medium effort cut tokens 26% for 4.2 ± 3.9 points of accuracy. A provider quietly routing you to a cheaper configuration looks, at first, like your model got terser — not dumber.

So log output tokens per question alongside scores, and alert on token drift too. It's the canary.

What you built

A drift watchdog with the livenerf shape: a frozen panel of borderline questions, a pinned harness, one run a day appended to a JSONL log, and a pre-registered statistical rule that turns "the model feels worse" into a yes-or-no answer. As of October 3, the real livenerf is 10 of 30 days in with zero missed days — baseline complete, first decision window open, first possible call around October 24. Your version starts its own clock today.

Honest limitations