Your coding agent bills by the token and writes like it has never heard of the bill. Every session starts polite — “heading in the right direction” — and ends with you paying for three pages of hedged prose around one diff. Worse, most of the tokens you pay for are not writing at all: they are reading. Your agent rereads the same logs, the same test output, the same half of the repo, turn after turn, and the meter runs on every reread.

There is a 109K-star project that treats this as a style problem. caveman (Apache-2.0, written in Go by JuliusBrussee) is built on one observation: your agent already knows how to be terse. It just needs rules that enforce it on the way out, and a proxy that shrinks everything on the way in. It is #2 on GitHub trending as of today, hit #1 on Hacker News back in April 2026, and — unusually for a viral project — its README leads with measured numbers, including the red ones.

This guide takes about twenty minutes. No account, no API key, no GPU.

Star history chart showing caveman's GitHub stars climbing past 90K within months of its April 2026 launch
caveman's star history, from the project's own README. It sits at 109K stars today. Source: JuliusBrussee/caveman (Apache-2.0).

What you'll need

  • Node.js 22.13+ (for the skill installer and the CLI)
  • One supported coding agent: Claude Code, Codex, Gemini CLI, Aider, Kilo, Qwen, opencode, Hermes, OpenClaw, or Pi for the proxy; the skill alone works in 30+ agents
  • An agent bill you can read — per-token, not per-request. If you are billed per request, skip to the limitations

Step 1 — Start with the small rock: the skill

Caveman comes in two sizes. Start small. The skill is a rule file that makes your agent answer in caveman: short sentences, no hedging, no filler. One command:

npx skills add JuliusBrussee/caveman -g

That is the whole install. Type /caveman if your agent does not wake up on its own. The skill only touches the agent's output — its mouth, not your prompts. That distinction is load-bearing: an Adobe Research paper that measured caveman-style compression (CAVEWOMAN, arXiv 2606.24083, eight models, five datasets) found it cuts realized output cost 1.4 to 2.4× per model, up to 3× in the best case — and that compressing the human's prompt into caveman-speak makes models answer longer and worse. Caveman never rewrites your side.

Step 2 — Find out where your tokens go

Before installing anything heavier, run the most useful five minutes in the README. Months of your agent history already sit on your disk; caveman learn reads it locally, read-only, and ranks the places your tokens go, biggest first, with a one-line fix behind each:

caveman learn             # Claude Code + Codex + Gemini CLI + opencode; aider via CAVEMAN_AIDER_ROOT

Then let it fix them, one at a time, only on your yes:

caveman learn implement   # hand the fixes to Claude Code or Codex, one diff at a time, applied only on your yes

implement re-measures after every change and undoes anything that did not make each message smaller. For fixes that cannot be re-counted — a new skill that only pays off when it gets used — caveman learn experiment runs it on for a stretch and off for a stretch over your own sessions, and gives no verdict before five sessions each way.

A caveman learn report: 41 sessions scanned, 855k tokens saved, token sinks ranked biggest first with suggested fixes
A caveman learn report from the project's README: 41 sessions scanned, token sinks ranked biggest-first, 855k tokens saved. Source: JuliusBrussee/caveman (Apache-2.0).

Step 3 — Graduate to the big rock: the proxy

The skill trims what the agent writes. The reading side is the bigger bill — JetBrains tested the skill alone on 86 real coding tasks (paired A/B, Claude Code 2.1.200) and found 8.5% fewer output tokens with no detectable quality change. Useful, but their finding was that an agent's bill is mostly reading, and no talking style fixes that. So the project built the thing that shrinks the reading: a proxy that runs on your machine, between your agent and the provider, compressing logs, test output, and diffs before every call.

npm install -g @caveman-ai/cli && caveman setup --install
caveman claude        # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · pi

No account, no API key. Your provider credentials stay where they were; the proxy sits in front of them. If you want the full installer — Claude Code hooks, the statusline badge, and detection of every supported agent on your machine — it is a pinned script instead:

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.sh | bash
irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.1.0/install.ps1 | iex

Step 4 — Measure it yourself

The maintainer's own advice outranks every number on the README: run the same task with and without caveman and compare your provider's billing page. That A/B beats all of the following, which are orientation, not promises:

  • Proxy benchmark (this repo, pinned): 54 runs of Claude Code across six cases, provider-reported input tokens, three runs per case, every answer checked against an exact oracle. Total: 885,793 input tokens direct vs 591,673 through caveman — −33.2%, with 18 of 18 answer checks passing.
  • Skill only (JetBrains, 86 real coding tasks): 8.5% fewer output tokens, about 10% of cost, no detectable quality change (sign test p = 0.82).
  • Output style (Adobe Research, CAVEWOMAN): 1.4–2.4× cost cut per model on output-side caveman style, up to 3× best case.

And the red row stays red, because the maintainer refuses to hide it: one proxy benchmark case — a dashboard HTML alert with no compression transform — came back +9.9%. No compression to apply, so caveman paid its own overhead and won nothing back. Compression has overhead. Sometimes it loses.

Step 5 — Use it in your own app: the middleware

Building an agent in code instead of running one in a terminal? Same shrinking, one wrapper around the call you already make. Apache-2.0 client, stable 1.0:

npm install @caveman-ai/middleware @caveman-ai/sdk        # TypeScript, plus your framework (ai, openai, …)
pip install 'caveman-middleware[langchain]' caveman-sdk   # Python 3.11+, swap the extra for your framework

Six lines of code and a local runtime, per the README’s walkthrough. It stacks with the skill and the proxy — most people start with the small rock and graduate.

What you built

A three-layer token diet for your coding agent: a skill that enforces terse output, a learn report that ranks your actual token sinks with fixes applied one diff at a time on your approval, and a local proxy that shrinks the read side before every provider call. Plus a measurement habit — the same-task A/B against your own billing page — that tells you whether it is winning on your workload.

Honest limitations

  • The skill's rules ride along as input tokens. About 1,000 estimated tokens per call for the full skill — on terse one-liner Q&A that can cost more than it saves. The repo's own eval: default caveman cut 3% of output tokens at the median against a plain “Answer concisely.” control, inside the noise. (ultracave: 35%.)
  • Never caveman your own prompts. The Adobe paper's other finding: compressing the human's prompt makes models answer longer and worse. The project only rewrites the agent's side — keep it that way.
  • Overhead is real. The benchmark's red row (+9.9% on an HTML case with nothing to compress) is the reminder: where there is no prose or repetition to cut, caveman pays its own tax.
  • Skip it entirely if you are billed per request rather than per token (GitHub Copilot premium requests: a shorter answer is the same request), or your workload is pure code generation with almost no prose to cut.
  • Measure yourself. If caveman loses on your workload, turn it off: npx -y github:JuliusBrussee/caveman -- --uninstall. The red row is the point — the day the project hides one is the day you stop trusting the green ones.

Talk like caveman. Measure like scientist. Keep the rocks that pay.