Watch a coding agent work for half an hour and you'll notice the real bottleneck isn't the model — it's the context window filling up with tool output. A Playwright snapshot: 56 KB. Twenty GitHub issues: 59 KB. One access log: 45 KB. After 30 minutes of routine agentic work, 40% of the context is gone, spent on bytes neither you nor the model ever needed to read.

Context Mode, by Mert Köseoğlu (mksglu/context-mode on GitHub), treats that output as data to be processed, not text to be pasted. It's an MCP server plus a set of agent hooks: code runs in an isolated sandbox and only the stdout reaches the model, documents get indexed into a local full-text search instead of being pasted wholesale, and session state survives compactions in a compact snapshot. When I looked it up on October 1, 2026, it was sitting at #3 on GitHub's daily trending list with 24,546 stars, a #1 spot on Hacker News with 570+ points, and an npm release at version 1.0.169. Its headline number is a 315 KB → 5.4 KB reduction on a real workflow — 98% of tool-output tokens kept out of context.

This tutorial installs Context Mode, runs every piece of it hands-on, and wires it into an agent. Every command below actually ran on my machine — version 1.0.169, verified end to end. One thing to know up front: the project is published under the Elastic License 2.0, which is source-available but not an OSI-approved open-source license. I'll say more about what that means in the caveats.

What you'll need

  • Node.js ≥ 22.5 or Bun — the package ships from npm.
  • An MCP-capable agent — Claude Code, Gemini CLI, VS Code or JetBrains Copilot, OpenCode, or KiloCode are all supported.
  • A few language runtimes — the sandbox supports twelve languages (JavaScript, TypeScript, Python, Shell, Ruby, Go, Rust, PHP, Perl, R, Elixir, C#), but only the ones installed on your machine count. My doctor run found 5 of 11 on this box; more on that in Step 2.
  • Ten minutes. The install is one npm command; the hooks configure themselves.

Honesty note on the headline number: 98% refers to tool-output tokens (315 KB → 5.4 KB on the project's measured workflow), not your entire context window. Your system prompt, conversation history, and the model's own reasoning still count. Context Mode shrinks the biggest controllable slice — everything the agent reads from tools — which is exactly the slice that grows fastest in long sessions.

Step 1 — Install Context Mode

The documented install is a single global npm command:

npm install -g context-mode

I installed it from the npm registry and pinned what I got:

npm install context-mode
npx context-mode --version
# 1.0.169

Version 1.0.169 is what every result below was verified against. The project moves fast — 195 releases since its February 2026 debut — so if a flag ever misbehaves, context-mode --help is authoritative. The CLI ships a handful of subcommands: doctor, upgrade, statusline, and hook (which wires the session hooks into your agent).

Step 2 — Let doctor grade your machine

Before touching the MCP tools, run the diagnostics:

context-mode doctor

Here's what it reported on my machine:

Runtime coverage: 5/11 languages
  JavaScript/TypeScript .... bun 1.4.2 (auto-detected)
  Python ................... python3 3.12.3
MCP server ................. PASS
Hooks ...................... configured

Two things worth reading carefully here. First, Bun was auto-detected for JavaScript and TypeScript — the sandbox prefers it for 3–5x faster JS/TS execution when it's present. Second, "5/11 languages" is a report on your machine, not a product limit: doctor can only sandbox languages whose runtimes are installed. If you live in Go or Rust, install those toolchains before expecting coverage. Everything else — MCP server, hooks — passed, so the wiring was ready.

Step 3 — Run your first sandboxed execution

Now the core idea. I started the MCP server over stdio and asked for its tool list. It introduced itself as context-mode, version 1.0.169, advertising all 11 tools: six sandbox tools (ctx_batch_execute, ctx_execute, ctx_execute_file, ctx_index, ctx_search, ctx_fetch_and_index) and five meta-tools (ctx_stats, ctx_doctor, ctx_upgrade, ctx_purge, ctx_insight). ctx_doctor — the in-agent twin of the CLI doctor — reported all checks passing.

Illustration of the sandbox: a glass cube containing tangled threads of raw data, emitting a single clean beam of output toward a robot observer
Credit: AI-generated illustration for AI Frontier Post

Then the first real execution — Python, computing factorials 0 through 7 and their sum:

{
  "tool": "ctx_execute",
  "arguments": {
    "language": "python",
    "code": "from math import factorial\nfacts = [factorial(i) for i in range(8)]\nprint(facts)\nprint('sum:', sum(facts))"
  }
}

What came back into context was only the stdout:

[1, 1, 2, 6, 24, 120, 720, 5040]
sum: 5914

That's the whole trick, and it's worth stating precisely: each ctx_execute call spawns an isolated subprocess with its own process boundary — scripts can't see each other's memory or state — and only the stdout enters the conversation context. The raw data never leaves the sandbox. The project's own tools table puts this at 56 KB → 299 B for a typical execution.

The README's canonical example makes the "before" concrete: 47 Read() calls dumping 700 KB to count lines across TypeScript files, versus one ctx_execute("javascript", …) call returning 3.6 KB. Same answer, two orders of magnitude less context.

Two behaviors kick in as outputs grow. When output exceeds 5 KB and you pass an intent, Context Mode switches to intent-driven filtering: it indexes the full output into the knowledge base, searches for the sections matching your intent, and returns only those — plus a vocabulary of searchable terms for follow-up queries. Authenticated CLIs (gh, aws, gcloud, kubectl, docker) work through credential passthrough, inheriting your environment and config paths without exposing them to the conversation.

Illustration of intent-driven filtering: a luminous funnel distilling scattered document icons and circuit traces into a single golden crystal
Credit: AI-generated illustration for AI Frontier Post

Step 4 — Build a searchable knowledge base

The sandbox keeps tool output out of context. The knowledge base keeps documents out of context. Instead of pasting files into the conversation, you index them once and search on demand. The engine is SQLite FTS5 with BM25 ranking, Porter stemming, trigram RRF fusion, headings weighted 5x, proximity reranking, fuzzy Levenshtein correction, and smart snippets.

I pointed ctx_index at two Markdown notes. It reported back: 2 files indexed, 2 sections. Then a ctx_search query returned BM25-ranked snippets with the source file paths attached — enough to answer from, without ever loading the documents into context. The project's measured number for this path is 60 KB → 40 B: index once, retrieve snippets forever.

A few operational details from the docs, all verified against the real README:

  • TTL cache: fetched content is cached under ~/.context-mode/content/ with a default 24-hour TTL (overridable per call with ttl: <ms>; ttl: 0 or force: true bypasses it), and stale entries are cleaned up after 14 days.
  • Parallel fetch: ctx_fetch_and_index accepts requests: [{url, source}, …] with concurrency: 1–8 for multi-URL ingestion in one call.
  • Batching: ctx_batch_execute runs multiple commands and searches multiple queries in a single call (measured 986 KB → 62 KB), with opt-in concurrency: 1–8 for I/O-bound batches.
  • Nuclear option: ctx_purge permanently deletes all indexed content from the knowledge base. It's the one destructive tool here — the schema even flags the others as read-only hints.

Step 5 — Watch the meter: ctx_stats

Skepticism is healthy with any "98% less" claim, so the project ships its own meter. ctx_stats reports context savings, call counts, and session statistics: bytes returned to context versus bytes sandboxed, bytes indexed, cache hits and cache bytes saved, tokens kept out, reduction percentage, tokens saved, and dollars saved — per session and lifetime, with a per-tool breakdown.

On my fresh scratch state it reported exactly what you'd expect from a new install: zeros across the board and $0.00 saved. The honest way to use it is as a before/after instrument: run ctx_stats at the start of a session, do your normal agentic work through the ctx_* tools, and read it again at the end. That's the only savings number that matters — yours, on your workload.

Step 6 — Wire it into your agent

The MCP tools are only half the system. The other half is hooks: PreToolUse, PostToolUse, UserPromptSubmit, PreCompact, SessionStart, and Stop. They route tool calls through the sandbox automatically, capture session events as they happen, build a resume snapshot before compaction, and restore state after it. The snapshot is priority-tiered to fit a 2 KB budget — critical state (active files, tasks, rules, decisions) is always preserved while lower-priority events drop first. Per-project SQLite keeps sessions isolated from each other.

The README quantifies the hooks' value: roughly 98% of tool-output tokens kept out of context with hooks, versus ~60% with instruction files (CLAUDE.md-style routing) alone. Instructions advise; hooks enforce.

Wiring it up, from the official docs:

Claude Code (recommended) — install as a plugin:

/plugin marketplace add mksglu/context-mode
/plugin install context-mode@context-mode

Then verify with /context-mode:ctx-doctor. The plugin registers all hooks and all 11 MCP tools, plus slash commands:

/context-mode:ctx-stats    per-tool savings, tokens, reduction ratio
/context-mode:ctx-doctor   runtimes, hooks, FTS5, versions
/context-mode:ctx-index    index a file or directory
/context-mode:ctx-search   search indexed content
/context-mode:ctx-upgrade  pull latest, rebuild, migrate cache, fix hooks
/context-mode:ctx-purge    delete all indexed content
/context-mode:ctx-insight  org analytics dashboard (context-mode.com/insight)

MCP-only (any agent):

claude mcp add context-mode -- npx -y context-mode

On platforms without slash-command support, you just type ctx stats, ctx doctor, ctx index, or ctx search in chat and the model calls the MCP tool. The SessionStart hook injects the routing instructions at runtime — nothing is written into your project. (One platform caveat from the docs: on Codex, PreToolUse routing currently supports deny rules only, pending upstream support for input rewrites.)

When to use Context Mode vs the alternatives

Context Mode isn't the only tool attacking context bloat. Here's how it sits next to approaches we've covered before:

  • RTK compresses shell output after the agent runs it — a proxy between agent and shell. Context Mode instead runs the command inside its own sandbox and never lets the raw output reach the agent at all. Use RTK when you want a drop-in output filter; use Context Mode when you want execution itself rerouted.
  • Prompt compression shrinks what you send. Context Mode shrinks what comes back. They stack — compress the prompt, sandbox the tools.
  • zvec-grep is semantic code search for your repo. Context Mode's FTS5 knowledge base is lexical (BM25 + stemming) but covers anything — docs, logs, fetched URLs — and lives behind the same MCP surface as the sandbox.
  • code-review-graph maps your repo's structure so the agent reads less code. Context Mode takes the complementary route: keep execution and retrieval behind the sandbox boundary entirely.
  • Instruction files alone (CLAUDE.md routing) get you ~60% of the way per the project's measurements; the hooks exist precisely because advice degrades under pressure and enforcement doesn't.

The honest summary: if your agents spend their sessions reading tool output, Context Mode attacks the largest line item. If your bloat is mostly prompt-side, start with compression instead.

Caveats, stated plainly

  • The license is Elastic License 2.0 (© 2026 Mert Koseoglu) — source-available, free to use, but not an OSI-approved open-source license. Read it before building commercial products or hosted services on top of this; the "open-source project" framing you'll see in trending lists is loose language.
  • Only stdout enters context. Exit codes, stderr, timing — if your script doesn't print it, the model never sees it. Write scripts that print what matters, and treat silent success as a real failure mode in agentic loops.
  • The 98% is tool-output tokens on the project's measured workflow, not a guarantee for yours. Run ctx_stats before and after a real session and believe your own numbers.
  • 5/11 languages on my machine was a mirror of installed runtimes, not a product cap. Install the toolchains you need, then re-run context-mode doctor.
  • Intent filtering needs both conditions: output over 5 KB and an explicit intent. Small outputs come back whole; that's by design.
  • ctx_purge is destructive — it permanently deletes the knowledge base. Everything else is read-only or additive.

The takeaway

Most context-window tooling treats the symptom — compressing, summarizing, or truncating text that's already in the window. Context Mode treats the cause: it moves execution and retrieval to the other side of a boundary the context window can't see, and only the distilled answer crosses over. The sandbox (ctx_execute), the knowledge base (ctx_index / ctx_search), and the hooks each do one job, and ctx_stats keeps all three honest.

Install it with npm install -g context-mode, run context-mode doctor, and let your next long agent session be the experiment. If the project's numbers hold on your workload, the most expensive tokens in your session — the ones you never read — simply stop being spent.