AI Frontier Post
Terminal code being funneled through a glowing compression funnel into a dense luminous cube
Compress before you send: Headroom squeezes agent context on your own machine — illustration generated for this article.

Your coding agent reads far more than it answers: tool outputs, log dumps, JSON blobs from MCP servers, whole files — and you pay for every token it reads. Prompt caching helps at the margins, but the payloads themselves are bloated: repeated keys, boilerplate, ceremony. Headroom is the bet that most of what an agent reads can be squeezed before it ever reaches the model — locally, reversibly, and without touching your code.

Headroom (headroomlabs-ai/headroom, ~74,500 GitHub stars at the time of writing, Apache-2.0) sits between your agent and the model and compresses everything the agent reads — tool outputs, logs, RAG chunks, files, conversation history. The project’s own seeded benchmarks claim 21–57% fewer input tokens depending on the scenario, answers unchanged, and compression runs entirely on your machine: no prompt or file content is sent anywhere to be compressed. The repo was pushed today and is under very active development. Here is the full hands-on.

What you’ll need

1. Install and health-check

pip install "headroom-ai[all]"
headroom doctor

doctor confirms routing works. The uv equivalent from the README is uv tool install --python 3.13 "headroom-ai[all]". That is the whole install: one package, and the headroom command is on your PATH.

2. Compress a real payload with the library

Start where the savings are biggest: repetitive tool output. This mirrors the README’s own inline example — build a realistic payload, compress it, and read off what you saved:

import json
from headroom import compress

# a realistic payload: 200 near-identical tool-result records,
# the kind of blob an MCP code-search call returns
records = [
    {"tool": "code_search", "file": f"src/mod_{i % 12}.py",
     "score": round(0.9 - i * 0.001, 4), "status": "ok",
     "snippet": "def handle(req):\n    return process(req)"}
    for i in range(200)
]
messages = [{"role": "user",
             "content": "Which modules handle requests? Results:\n"
                        + json.dumps(records)}]

result = compress(messages, model="gpt-4o")
print(f"saved {result.tokens_saved} tokens "
      f"({result.compression_ratio:.0%} smaller)")
# send result.messages to the model exactly as usual

The returned result.messages is a drop-in for the original list — the README’s example passes it straight to the OpenAI client. For the number that applies to your traffic rather than a synthetic blob, the README’s advice is blunt: run headroom savings against your own sessions.

What happened under the hood: a ContentRouter detects the content type and picks a compressor — SmartCrusher for JSON, CodeCompressor for source (AST-aware across Python, JS/TS, Go, Rust, Java, C/C++, C#, and PHP), and Kompress-v2-base (the project’s own HuggingFace model) for prose. A CacheAligner flags volatile content that would bust the provider’s KV-cache prefix, and only fresh bytes are compressed — the frozen prefix stays byte-identical so the provider cache survives. Originals are cached locally and retrievable on demand, so compression is reversible.

A messy stack of JSON documents and log lines routed through a glowing node into a single dense packet
Route, then crush: the ContentRouter picks SmartCrusher for JSON, CodeCompressor for code, or the Kompress model for prose — illustration generated for this article.

3. Reproduce their proof numbers

The benchmark table in the README is seeded and offline, which means you can get the project’s exact numbers yourself:

git clone https://github.com/headroomlabs-ai/headroom && cd headroom
uv run python benchmarks/index_proof_table.py --seed 20260902
ScenarioBeforeAfterSaved
Code search (100 results)17,19913,59721%
SRE incident debugging55,95724,34057%
Codebase exploration58,80133,89542%
GitHub issue triage46,06732,42930%

Savings scale with repetitiveness: repeated JSON arrays and log lines clear 90% in their latency bench; prose and already-dense output compress very little. Compression itself costs 0.21 ms p50 on a 10K-token JSON result — the README’s line is that it “does not show up in agent latency.” On accuracy, their tier-1 evals report GSM8K 0.870 → 0.870, no detectable difference on TruthfulQA at N=100, 97% on SQuAD v2 at 19% compression, and 97% on BFCL tool-calling at 32% compression.

4. Zero-code-change proxy

Don’t want to touch your app? Run the proxy and point any OpenAI-compatible client at it:

headroom proxy --port 8787

Set your client’s base URL to http://localhost:8787/v1 and use it normally — traffic compresses on the way through. Then watch it live:

headroom dashboard   # live savings (proxy must be running)
headroom perf

The README’s claim is that any OpenAI-compatible client works through the proxy, in any language — the compression layer speaks the API your tools already speak.

5. Wrap your coding agent in one command

headroom wrap claude     # claude | codex | grok | copilot | cursor | aider |
                         # opencode | cline | continue | goose | openhands |
                         # openclaw | vibe | omp | zcode
headroom unwrap claude   # undo

Wrapping starts the local proxy and launches the agent routed through it. For Claude Code it also installs Serena for semantic code navigation, registered as a local-scope MCP server for that project only. Prefer MCP-native wiring instead? headroom mcp install hands any MCP client three tools — headroom_compress, headroom_retrieve, headroom_stats — and there is a shared cross-agent memory store (Claude, Codex, Gemini, Grok) with automatic dedup if you run several agents.

An AI coding agent beside a monitor showing a chat window and a low token-usage gauge
One command in, one command out: wrapping routes your agent through the local proxy — illustration generated for this article.

6. Shrink the output side too

Everything so far shrinks the prompt you send. You also pay for every token the model writes back — and much of it is ceremony: “Great, let me…” preambles, code re-printed back at you, deep reasoning spent on routine steps. Headroom trims it from the proxy, with no code change:

export HEADROOM_OUTPUT_SHAPER=1   # off by default
headroom proxy --port 8787
headroom output-savings
# Reduction: 31.7%  (95% CI 27.7% … 35.7%)   [estimated]

Two mechanisms: verbosity steering (a short “be terse” note appended to the end of the system prompt, so your prompt cache still hits) and effort routing (thinking effort dialed down when a turn is just the model resuming after a tool result). Output savings are counterfactual — nobody sees what the model would have written — so the number is reported as an estimate with a confidence band. For a measured number instead, hold out 10% of conversations as an unshaped control with HEADROOM_OUTPUT_HOLDOUT=0.1 and the dashboard reads measured instead of estimated.

What you built

A compression layer around your agent stack that you never have to think about again: repetitive tool output and logs get squeezed before they reach the model, the proxy or wrapper applies it with zero code changes, the output shaper trims the model’s own ceremony, and every byte of the original stays retrievable locally if the model ever needs it. On the project’s seeded benchmarks that is a fifth to well over half of input tokens back — on your traffic, headroom savings tells you the real number.

Honest limitations

Related articles

Tutorials

Give your AI agent a real memory: hands-on with Hindsight, the local-first memory layer

September 29, 2026 · 13 min read
Tutorials

Stop letting your coding agent grep: hands-on with CodeGraph, the 72K-star pre-indexed code knowledge graph

October 1, 2026 · 12 min read
Tutorials

Stop burning tokens on grep: find code by asking what it does — hands-on with jevgrep

October 3, 2026 · 8 min read