RTK: cut your AI agent's token burn with a Rust CLI output proxy
Every time your coding agent runs a shell command, the full output lands in its context window: the 40-line git status, the 3,000-line test log, the directory listing with every timestamp and permission flag. None of it was designed for a model to read, and all of it costs you input tokens. RTK — the Rust Token Killer — sits between your agent and the shell and compresses that output before the model ever sees it. It is a single binary with zero dependencies, it covers 100+ dev commands, and on this week's runs I measured 41–74% smaller outputs at ~10ms of overhead. Here is how to install it, measure it honestly, and wire it into your agent so every shell call gets cheaper.

Why this is blowing up now
RTK (rtk-ai/rtk, Apache-2.0) started 2026-01-22 and now sits at 81,838 stars and 5,190 forks, with commits landing daily and releases shipping at a relentless pace — v0.50.0 went out on 2026-09-24, and the version string in the README already lags behind it. It ranks on GitHub's trending AI lists this week alongside agent harnesses and memory layers, because it attacks the most boring line item on every agent bill: the noise in shell output.
The insight is simple. Tools like git diff, ls, pytest, and npm test were designed for human terminal readability: progress spinners, ASCII headers, repeated success lines, decorative separators. Fed straight into an LLM, that decoration becomes thousands of input tokens that carry zero signal. RTK applies four deterministic strategies per command — smart filtering, grouping, truncation, and deduplication — and hands the agent only what it can act on.

What you'll need
- A Linux, macOS, or Windows machine — RTK ships pre-built binaries for all three.
- A terminal and a git repo to experiment on. No account, no API key, no GPU, no model downloads. The binary is a few megabytes; nothing phones home by default (telemetry is opt-in — verified below).
- About 20 minutes. If you use Claude Code, Cursor, or a similar agentic coding tool, you will also wire in a hook in the final step — but every measurement in this tutorial works without any agent at all.
Step 1 — Install the binary
Pick your platform. These are the project's own documented paths; I used the Linux quick-install for this tutorial and it dropped a working rtk 0.50.0 into ~/.local/bin in seconds.
macOS (Homebrew, recommended by the project):
brew install rtk
Windows (winget):
winget install rtk-ai.rtk
Linux / macOS (quick install):
curl -fsSL https://raw.githubusercontent.com/rtk-ai/rtk/refs/heads/master/install.sh | sh
# installs to ~/.local/bin; add it to PATH if needed:
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
Prefer not to pipe a script to your shell? Grab the tarball for your platform from the releases page (rtk-x86_64-unknown-linux-musl.tar.gz on Linux, .zip on Windows) and drop the single binary on your PATH. Rust developers can cargo install --git https://github.com/rtk-ai/rtk.
Verify the install — and watch out for one gotcha:
rtk --version # expect: rtk 0.50.0 (or newer; the README's 0.28.2 lags reality)
rtk gain # expect: the savings dashboard, or "No tracking data yet."
The gotcha: a different project also named rtk (Rust Type Kit) exists on crates.io. If rtk gain fails or prints something unrelated, you have the wrong binary — install from the GitHub releases instead. rtk gain working is the reliable identity check.
Step 2 — Measure the savings yourself, honestly
Before trusting any percentage, measure it. Build a scratch repo with enough files to make the noise visible, then compare raw output against RTK's output byte-for-byte:
mkdir -p rtk-demo/src && cd rtk-demo
git init -q
for i in $(seq 1 25); do echo "# module $i" > src/mod_$i.py; done
git add -A && git -c [email protected] -c user.name=you commit -qm init
# make some edits so status/diff have something to chew on
for i in 2 3 4 5 6 7; do echo "# edit $i" >> src/mod_$i.py; done
Now the comparison. I ran this on a Linux VM; your absolute numbers will differ, but the pattern holds:
git status | wc -c # raw
rtk git status | wc -c # compact
Here is what I measured:
| Command pair | Raw output | Via RTK | Reduction |
|---|---|---|---|
git status | 529 B | 153 B | 71% |
ls -la src → rtk ls src | 1,417 B | 366 B | 74% |
git log | 116 B | 34 B | 71% |
git diff | 870 B | 511 B | 41% |
Two things worth understanding before you quote those numbers. First, RTK measures reductions in bash output, not in your total bill — the README is explicit about this, and it is right: command output is one contributor to input tokens alongside the system prompt, the conversation, and everything else the agent reads. A 70% smaller git status is real savings; it is not a 70% smaller invoice.
Second, RTK estimates tokens as bytes ÷ 4 — it ships no tokenizer. Treat the percentages as reliable and the absolute token counts as approximate. The project's own docs say exactly this, and the dashboard you will see in Step 5 labels them as estimates.
Step 3 — Learn the daily verbs
RTK is a proxy: you prefix the command you would have run, and it runs it, filters the output, and prints the compact version. The coverage is broad — the built-in command list spans files, git, GitHub/GitLab CLIs, test runners, linters, package managers, Docker, kubectl, AWS, and databases. The verbs you will reach for most:
rtk ls src # compact tree: name + size, no permission/timestamp noise
rtk tree . # project structure, token-optimized
rtk read src/mod_2.py # smart file reading: structure over full bodies
rtk read src/mod_2.py -l aggressive # signatures only, bodies stripped
rtk grep -r "edit" src # matches grouped by file, long lines truncated
rtk diff a.py b.py # condensed diff (exit codes: 0 identical, 1 different, 2 read error)
Git gets the most polish:
rtk git status # compact stat format, grouped by state
rtk git log # hash, author, subject only
rtk git diff # condensed diff, context trimmed, headers stripped
rtk git add . # -> "ok"
rtk git commit -m "msg" # -> "ok abc1234"
rtk git push # -> "ok main"
One honest behavioral note: RTK is a faithful proxy, which means it inherits native quirks. rtk grep "edit" . fails with "Is a directory" exactly like grep "edit" . does — pass -r, or use rtk rg if you have ripgrep. It does not try to outsmart the tool; it shrinks what the tool prints.

Step 4 — Collapse your test output
Test logs are the worst offender in agent workflows: a green suite prints hundreds of identical PASSED lines, and the one failure drowns in them. RTK's answer is failures-only output. Set up a tiny suite with one failure:
mkdir -p tests
printf 'def test_a(): assert 1 + 1 == 2\n' > tests/test_a.py
printf 'def test_b(): assert 1 + 1 == 3\n' > tests/test_b.py
python -m pytest -q # raw: full traceback + summary
rtk test python -m pytest -q # failures only
Raw output: 528 bytes of dots, tracebacks, and summary. RTK's version:
[FAIL] FAILURES:
FAILED tests/test_b.py::test_b - assert (1 + 1) == 3
SUMMARY:
1 failed, 1 passed in 0.05s
[full output: rtk recall 00f80e519e92]
152 bytes — a 71% reduction — and nothing lost that matters: the failing test, the assertion, and the summary are all there. And that last line is the escape hatch: RTK stores the full output locally, so when you need the complete traceback you recover it by content hash:
rtk recall 00f80e519e92 # prints the full original output
I verified rtk recall restores the complete unfiltered log. This is the design decision that makes aggressive filtering safe: the compact output is the default, the full output is one command away. rtk test works with any command, not just pytest — the same pattern covers rtk pytest, rtk jest, rtk vitest, rtk cargo test, and rtk go test when those runners are on your PATH.
Step 5 — Read the savings dashboard
RTK tracks every proxied command in a local SQLite database and reports cumulative savings. This is where the "up to 90%" marketing claim becomes something you can audit on your own machine:
rtk gain
RTK Token Savings (Global Scope)
════════════════════════════════════════════════════════════
Total commands: 15
Input tokens: 1.8K
Output tokens: 1.0K
Tokens saved: 763 (43.1%)
Total exec time: 157ms (avg 10ms)
Efficiency meter: ██████████░░░░░░░░░░░░░░ 43.1%
By Command
───────────────────────────────────────────────────────────────────────
# Command Count Saved Total% Time Impact
───────────────────────────────────────────────────────────────────────
1. rtk ls src 2 430 70.0% 18ms ██████████
2. rtk git status 2 150 74.3% 18ms ███░░░░░░░
3. rtk git diff 1 129 50.2% 13ms ███░░░░░░░
...
That is my real dashboard after this tutorial's measurements: 763 estimated tokens saved across 15 commands, at an average of 10ms overhead per call — the project's "<10ms overhead" claim held up. rtk gain --history shows the same table; rtk discover goes further and scans your Claude Code history for commands you ran raw that RTK could have compressed.
Privacy, since this tool watches your shell: tracking is local-only, and telemetry is opt-in. rtk telemetry status on a fresh install reports consent: never asked / enabled: no. Configuration lives in a plain TOML file (~/.config/rtk/config.toml) where you can tune ignored directories (.git, node_modules, __pycache__ are excluded by default) and display options via rtk config.
Step 6 — Wire it into your agent
Prefixing commands manually is fine for measurement; the real payoff is the hook. rtk init registers RTK with your coding agent so Bash calls get rewritten automatically — git status becomes rtk git status before it executes, and the agent receives compact output without ever typing rtk itself:
rtk init -g # Claude Code (global)
rtk init -g --gemini # Gemini CLI
rtk init -g --codex # Codex
rtk init --agent cursor # Cursor
# also: windsurf, cline, kilocode, antigravity, kimi, pi, hermes,
# droid, vibe, omp, trae, opencode
I ran rtk init -g against a sandboxed home directory to see exactly what it does. It writes three things and then asks you to finish the wiring:
~/.claude/RTK.md— an awareness doc telling the agent its command output is condensed, to treat it as complete, and to re-run viartk proxy <cmd>only when a result is unusable (empty when output was expected, contradicting its exit code, or garbled).- A
@RTK.mdreference appended to~/.claude/CLAUDE.md. - A PreToolUse hook — it prints the JSON to paste into
~/.claude/settings.json(in interactive mode it offers to patch the file for you):{ "hooks": { "PreToolUse": [{ "matcher": "Bash", "hooks": [{ "type": "command", "command": "rtk hook claude" }] }]} }
Then restart the agent and run git status — the hook rewrites it transparently. Two practical notes from testing: create ~/.claude first if it does not exist (rtk init errors out otherwise), and know the hook's boundary — it only fires on Bash tool calls. Claude Code's built-in Read, Grep, and Glob tools do not pass through the Bash hook, so they are not auto-rewritten; for those workflows, call rtk read, rtk grep, or rtk find explicitly, or just use shell commands instead of the built-ins.
When to use RTK vs the alternatives
vs ccusage / ccburn: those tools monitor your spend; RTK reduces the input side of it. They answer different questions ("where did my money go" vs "how do I spend less on shell output"), and they compose well — several Claude Code config guides recommend running both.
vs Headroom and context compressors: same problem space, different layer. Compressors summarize everything the agent reads — tool output, logs, conversation history, RAG chunks — often with a local model. RTK is deterministic, rule-based, and scoped to the shell boundary: no model, no latency beyond ~10ms, no subscription. Use RTK when the bloat is command output; reach for a compressor when the bloat is prose, history, or files the agent read directly.
vs prompt caching: orthogonal and complementary. Smaller tool outputs mean cheaper cache writes and reads, and the project's docs note RTK does not break prompt caching — filtered output is stored in history and cached normally.
vs doing nothing: if your agent sessions are short and your commands are small, the absolute savings are small too — RTK's own rtk read on a two-line file saved exactly zero bytes in my tests, correctly. The value compounds in long agentic sessions where the same noisy commands run dozens of times.
Caveats, stated plainly
- The headline percentages are reductions in bash output, not in your total API bill. Command output is one input-token contributor among several; the README says this explicitly, and any tutorial that implies otherwise is misreading the tool.
- Absolute token numbers are
bytes ÷ 4estimates. The percentages are the reliable part. - Leave interactive commands alone — SSH sessions, REPLs, editors, and anything that needs a TTY. The filters are built for batch output.
- Make sure you have this rtk: the crates.io
rtk(Rust Type Kit) is an unrelated project.rtk gainprinting the savings dashboard is the identity check. - The project moves fast — 371 releases at last count, and the README's version string already lagged the binary I installed (0.28.2 documented vs 0.50.0 installed). Pin or check
rtk --versionif a flag in this tutorial ever misbehaves;rtk --helpis authoritative.
The takeaway
RTK does one thing — shrink the shell output your agent reads — and does it with a single dependency-free binary, per-command filters tuned for the dev tools you actually run, a recall hatch for the full output, and an honest dashboard that shows your own numbers instead of a marketing claim. Install it in minutes, measure it against your own repos the way this tutorial did, and if the percentages hold, wire in the hook and stop paying to transmit ASCII art to a model. That is the entire pitch, and it survives contact with a real terminal.