
Your coding agent reads far more than it answers: tool outputs, log dumps, JSON blobs from MCP servers, whole files — and you pay for every token it reads. Prompt caching helps at the margins, but the payloads themselves are bloated: repeated keys, boilerplate, ceremony. Headroom is the bet that most of what an agent reads can be squeezed before it ever reaches the model — locally, reversibly, and without touching your code.
Headroom (headroomlabs-ai/headroom, ~74,500 GitHub stars at the time of writing, Apache-2.0) sits between your agent and the model and compresses everything the agent reads — tool outputs, logs, RAG chunks, files, conversation history. The project’s own seeded benchmarks claim 21–57% fewer input tokens depending on the scenario, answers unchanged, and compression runs entirely on your machine: no prompt or file content is sent anywhere to be compressed. The repo was pushed today and is under very active development. Here is the full hands-on.
headroom-ai package is the TypeScript SDK only and provides no headroom command.pip install "headroom-ai[all]"
headroom doctor
doctor confirms routing works. The uv equivalent from the README is uv tool install --python 3.13 "headroom-ai[all]". That is the whole install: one package, and the headroom command is on your PATH.
Start where the savings are biggest: repetitive tool output. This mirrors the README’s own inline example — build a realistic payload, compress it, and read off what you saved:
import json
from headroom import compress
# a realistic payload: 200 near-identical tool-result records,
# the kind of blob an MCP code-search call returns
records = [
{"tool": "code_search", "file": f"src/mod_{i % 12}.py",
"score": round(0.9 - i * 0.001, 4), "status": "ok",
"snippet": "def handle(req):\n return process(req)"}
for i in range(200)
]
messages = [{"role": "user",
"content": "Which modules handle requests? Results:\n"
+ json.dumps(records)}]
result = compress(messages, model="gpt-4o")
print(f"saved {result.tokens_saved} tokens "
f"({result.compression_ratio:.0%} smaller)")
# send result.messages to the model exactly as usual
The returned result.messages is a drop-in for the original list — the README’s example passes it straight to the OpenAI client. For the number that applies to your traffic rather than a synthetic blob, the README’s advice is blunt: run headroom savings against your own sessions.
What happened under the hood: a ContentRouter detects the content type and picks a compressor — SmartCrusher for JSON, CodeCompressor for source (AST-aware across Python, JS/TS, Go, Rust, Java, C/C++, C#, and PHP), and Kompress-v2-base (the project’s own HuggingFace model) for prose. A CacheAligner flags volatile content that would bust the provider’s KV-cache prefix, and only fresh bytes are compressed — the frozen prefix stays byte-identical so the provider cache survives. Originals are cached locally and retrievable on demand, so compression is reversible.

The benchmark table in the README is seeded and offline, which means you can get the project’s exact numbers yourself:
git clone https://github.com/headroomlabs-ai/headroom && cd headroom
uv run python benchmarks/index_proof_table.py --seed 20260902
| Scenario | Before | After | Saved |
|---|---|---|---|
| Code search (100 results) | 17,199 | 13,597 | 21% |
| SRE incident debugging | 55,957 | 24,340 | 57% |
| Codebase exploration | 58,801 | 33,895 | 42% |
| GitHub issue triage | 46,067 | 32,429 | 30% |
Savings scale with repetitiveness: repeated JSON arrays and log lines clear 90% in their latency bench; prose and already-dense output compress very little. Compression itself costs 0.21 ms p50 on a 10K-token JSON result — the README’s line is that it “does not show up in agent latency.” On accuracy, their tier-1 evals report GSM8K 0.870 → 0.870, no detectable difference on TruthfulQA at N=100, 97% on SQuAD v2 at 19% compression, and 97% on BFCL tool-calling at 32% compression.
Don’t want to touch your app? Run the proxy and point any OpenAI-compatible client at it:
headroom proxy --port 8787
Set your client’s base URL to http://localhost:8787/v1 and use it normally — traffic compresses on the way through. Then watch it live:
headroom dashboard # live savings (proxy must be running)
headroom perf
The README’s claim is that any OpenAI-compatible client works through the proxy, in any language — the compression layer speaks the API your tools already speak.
headroom wrap claude # claude | codex | grok | copilot | cursor | aider |
# opencode | cline | continue | goose | openhands |
# openclaw | vibe | omp | zcode
headroom unwrap claude # undo
Wrapping starts the local proxy and launches the agent routed through it. For Claude Code it also installs Serena for semantic code navigation, registered as a local-scope MCP server for that project only. Prefer MCP-native wiring instead? headroom mcp install hands any MCP client three tools — headroom_compress, headroom_retrieve, headroom_stats — and there is a shared cross-agent memory store (Claude, Codex, Gemini, Grok) with automatic dedup if you run several agents.

Everything so far shrinks the prompt you send. You also pay for every token the model writes back — and much of it is ceremony: “Great, let me…” preambles, code re-printed back at you, deep reasoning spent on routine steps. Headroom trims it from the proxy, with no code change:
export HEADROOM_OUTPUT_SHAPER=1 # off by default
headroom proxy --port 8787
headroom output-savings
# Reduction: 31.7% (95% CI 27.7% … 35.7%) [estimated]
Two mechanisms: verbosity steering (a short “be terse” note appended to the end of the system prompt, so your prompt cache still hits) and effort routing (thinking effort dialed down when a turn is just the model resuming after a tool result). Output savings are counterfactual — nobody sees what the model would have written — so the number is reported as an estimate with a confidence band. For a measured number instead, hold out 10% of conversations as an unshaped control with HEADROOM_OUTPUT_HOLDOUT=0.1 and the dashboard reads measured instead of estimated.
A compression layer around your agent stack that you never have to think about again: repetitive tool output and logs get squeezed before they reach the model, the proxy or wrapper applies it with zero code changes, the output shaper trims the model’s own ceremony, and every byte of the original stays retrievable locally if the model ever needs it. On the project’s seeded benchmarks that is a fifth to well over half of input tokens back — on your traffic, headroom savings tells you the real number.
min_input_words come back byte-identical — that is the README’s own “when to skip” section talking.headroom savings.