AI Frontier Post
A glowing speedometer measuring tokens per second above a GPU circuit board
Measure what an agent actually feels: agentperf-local replays real agent traffic against your local model server — illustration generated for this article.

Everyone quotes a tokens-per-second number for their local model, and almost none of them mean anything for agents. A bare completion benchmark sends one short prompt and reads the tokens. An agent sends the whole conversation so far — every tool result, every file it opened — on every turn. agentperf-local is the open-source tool from Artificial Analysis that benchmarks how fast your machine serves an agent, not a prompt: it replays recorded agent conversations against your model server and reports throughput and latency per turn.

agentperf-local (ArtificialAnalysis/aa-agentperf-local, ~85 GitHub stars at the time of writing, Apache-2.0) works two ways: point it at a server you already run (llama.cpp, LM Studio, vLLM, SGLang, Ollama), or let it download a pinned model, start the server itself, and run the replay. The default replay is eight recorded agent tasks with 168 model turns, each turn carrying the full conversation so far — the shape real agentic traffic actually takes. Here is the full hands-on.

What you’ll need

1. Install and see your options

uv tool install agentperf-local
agentperf-local doctor
agentperf-local deployment-options --profile-id qwen38-27b-q4-k-m

doctor prints local hardware facts without identifiers — a sanity check that your GPU and RAM are what you think they are. deployment-options shows which frameworks can serve one catalog recipe on this machine (the default profile is qwen38-27b-q4-k-m; pass --profile-id for another). macOS, Linux, and Windows are supported — with the caveats the README is explicit about: on Windows, managed runs need an NVIDIA GPU and llama.cpp, and vLLM and SGLang run only on Linux.

2. The six-turn quick check

Before you commit to a full run, prove the plumbing works with the synthetic mini replay — aa-mini-v1, six turns, needs only 8,192 tokens of context:

agentperf-local managed-run \
  --profile-id gemma4-12b-it-q4-0 \
  --framework llama-cpp \
  --replay aa-mini-v1 \
  --context-tokens 8192 \
  --output-dir results/gemma4-12b-mini

managed-run downloads the pinned model into the Hugging Face cache, checks every file’s SHA-256, starts the server on localhost, runs the replay, and stops the server. The tool never installs serving frameworks for you — each recipe pins the exact version it expects. Prefer the guided path? agentperf-local tui --replay aa-mini-v1 opens the full-screen app, where you choose a smaller context on the Setup screen (arrows move, Enter continues, Escape goes back, ? opens help, q quits). Two honest labels to carry forward: a run below 65,536 tokens is marked reduced: true, and the mini replay’s results are explicitly not comparable with anything.

3. The full 168-turn run

The default replay, agentperf-default-v1, is eight recorded agent tasks with 168 model turns, and it uses the exact output policy — every turn generates its recorded number of tokens, which is what makes results comparable across machines and frameworks:

agentperf-local managed-run \
  --profile-id qwen38-27b-q4-k-m \
  --framework llama-cpp \
  --output-dir results/qwen38-27b
A dark terminal window showing a bar chart of LLM serving benchmark results
The number that matters is throughput under agent-shaped traffic — full conversations, every turn — illustration generated for this article.

Why agent-shaped traffic changes the answer: each request carries the conversation so far, so later turns are long-context prefills, not short prompts. A server that looks fast on single-shot prompts can fall over when the 120th turn of an agent session arrives with 50,000 tokens of history. That is the difference this tool measures. On Apple Silicon, the qwen38-27b-splash-dflash recipe serves on Splash, Inco AI’s Metal engine (M3 or newer, macOS 26.4 or later; install with brew install incoai/tap/splash) — the recipe pins Splash 1.2.1 and the model package by commit, so the server never silently resolves a newer revision.

4. Benchmark the server you already run

Already serving with llama.cpp, LM Studio, vLLM, or SGLang? Point the replay at it instead:

agentperf-local run \
  --base-url http://127.0.0.1:8080/v1 \
  --model served-model \
  --output-dir results/my-server
ServerTypical base URL
llama.cpphttp://127.0.0.1:8080/v1
LM Studiohttp://127.0.0.1:1234/v1
vLLMhttp://127.0.0.1:8000/v1
SGLanghttp://127.0.0.1:30000/v1

Use the model name your server reports. Two gotchas the README is explicit about. First, the exact policy sends ignore_eos, which is not part of the OpenAI API — the tool checks that your server honors it before the replay, and if it doesn’t, it stops and names the flag to change. Second, Ollama cannot honor ignore_eos: the tool detects Ollama, warns you, and falls back to the recorded policy, which lets the model stop on its own and reports end-to-end latency as a normalized estimate. That result is not directly comparable with exact runs — don’t benchmark Ollama against a llama.cpp run and read the winner as truth. If your server needs an API key, put it in an environment variable and pass the variable’s name with --api-key-env — no flag takes a literal key.

5. Add real tool time (optional)

By default the replay skips the time between turns. Real agents spend wall-clock time running tools. Two modes add it back:

agentperf-local run \
  --base-url http://127.0.0.1:8080/v1 \
  --model served-model \
  --output-dir results/live-tools \
  --tool-mode fixed_delay

fixed_delay sleeps the recorded tool time after each turn. --tool-mode live goes further and actually runs the recorded shell commands inside Docker containers, so tool time is measured, not slept. It is opt-in for a reason: it needs Docker plus prebuilt images (scripts/build-default-containers.sh, and scripts/install-swebench-validation.sh once on arm64 hosts; they are bash scripts, so Windows users run them from Git Bash or WSL). The README’s security caveats are worth reading in full — the recorded commands are untrusted input, containers have no network by default (a replay manifest can ask for one; --live-network none forces none), and the task workspace is mounted read-write under your output directory.

A stopwatch hovering over a small AI agent running on a glowing circuit path
Time what the agent actually waits for: generation, prefill, and tool round-trips — illustration generated for this article.

6. Read the results

Each run writes a new directory:

results/qwen38-27b/
├── turns.jsonl      # per-turn measurements
├── tasks.json
├── tools.json
├── failures.json
├── summary.json     # the headline numbers
└── measurement.json

summary.json is written last — a crashed run has no summary. It holds the headline numbers: output tokens per second, and median and p95 first-token and turn times. Managed runs also write the server log and a deployment record, so you can prove what was actually served. A note on privacy: the results are private, but they can contain paths, model labels, and errors, and a run sends the recorded prompts to the model server — point the base URL at a remote server and those prompts leave your machine.

What you built

A repeatable measurement of how your machine serves agents: an installed benchmark CLI, a quick-check config that proves your setup, a full 168-turn comparable run against a pinned model, and a results directory with tokens-per-second and latency percentiles you can diff against the next framework, the next quantization, or the next GPU. Submitting results to Artificial Analysis is optional — nothing is uploaded unless you ask (prepare-submission builds the file, submit sends it, submission-status reads it back).

Honest limitations

Related articles

Tutorials

Cut your AI agent’s token bill by 20–60%: hands-on with Headroom, the local compression layer

October 7, 2026 · 8 min read
Tutorials

Run a 125-billion-parameter model on your own gaming PC: hands-on with Strata

October 4, 2026 · 8 min read
Tutorials

Run Claude Code on DeepSeek: route every coding agent through one local gateway

October 5, 2026 · 8 min read