AI Frontier Post
A stylized magpie perched above a glowing network of AI model nodes
Every agent's model, one place - magpie routes any coding agent to any model through one local gateway. Illustration: AI Frontier Post.

Every coding agent keeps its model in its own file, in its own format, with its own keys and base URLs. Your Claude Code subscription works in Claude Code and nowhere else. Your cheap DeepSeek key works in the agent that lets you paste a base URL, and nowhere else. When a quota runs out at 3 pm, you start editing config files - and every agent has a different idea of which vendors it even supports.

Magpie puts all of it in one place. It is a local model gateway - one process, one screen - that runs Claude Code on Kimi, Codex on DeepSeek, Gemini CLI on GLM, or OpenCode on your ChatGPT plan. It ships a gateway on 127.0.0.1:3425 that speaks the OpenAI Chat, OpenAI Responses, Anthropic Messages, and Gemini protocols, translating between them with streaming, tool calls, and reasoning intact. Pick any model for any agent from a single list; magpie rewrites only that one key in the agent's own config file, atomically, leaving your comments and formatting alone. It is MIT-licensed, written in Go, runs on macOS, Windows, and Linux, and picked up 5,100+ stars in about 12 days.

What you get at the end: every agent on your machine pointed at whatever model you want, plus routing groups that fail over to the next model when a quota runs dry - so your agent never sees the error. Here is the whole thing, hands-on.

What you'll need

1. Install magpie

The one-liner from the README:

curl -fsSL https://usemagpie.ai/install.sh | sh

Alternatives: download the app from usemagpie.ai (Mac builds are signed and notarised, and every build updates itself), go install github.com/yetone/magpie@latest, or the Docker image ghcr.io/yetone/magpie. Behind a firewall, the installer takes --proxy or --mirror.

2. Add your first provider

Open magpie and go to Providers -> Add provider, pick a preset, paste a key - done. Or do it from the terminal:

magpie provider add deepseek sk-...

Presets include Anthropic, OpenAI, Google Gemini, DeepSeek, Kimi, Zhipu GLM, MiniMax, StepFun, Qwen, Mistral, Groq, xAI, OpenRouter, Together, Fireworks, Ollama, LM Studio - plus any OpenAI-compatible or Anthropic-compatible URL. The model list comes from the vendor itself, with names and reasoning levels filled in from models.dev, so a model released this morning shows up on the next refresh. No model list is baked into magpie. Already configured elsewhere? Import reads the providers you set up in CC Switch, Claude Code, Codex, and Alma.

3. Point an agent at any model

On the Agents page, click a model value and a filtered list opens holding every model of every provider as provider/model. Pick one - magpie rewrites just that key in the agent's own config file, atomically. The CLI form:

magpie claude deepseek/deepseek-v4-pro   # Claude Code on DeepSeek
magpie codex moonshot/kimi-k2.5          # Codex on Kimi K2.5

Start a new agent session and it uses the new model. The same list and the same click work across all 35+ supported agents - Claude Code, Claude Desktop, Codex, Gemini CLI, OpenCode, Cursor CLI, Cline, Zed, VS Code Chat, Copilot CLI, and the rest.

A glowing model picker dropdown over a dark terminal
One list of every model, every provider - pick and the agent's config is rewritten atomically. Illustration: AI Frontier Post.

4. Share your subscriptions with every agent

This is the part that surprised me. An agent you are already signed in to becomes a provider in magpie - your Claude, ChatGPT, Copilot, Gemini (Code Assist), or Grok sign-in works in every other agent, with no key to copy. Several accounts per subscription are supported, with failover between them:

5. Build a routing group that survives quota exhaustion

A routing group is several models that an agent picks as one. When one hits a rate limit or runs out of quota, the next one answers - your agent never sees the error:

magpie group add "Opus anywhere" models=claude/claude-opus-5-5,copilot/claude-opus-5.5 routing=smart
magpie claude group/opus-anywhere   # fails over between subscriptions

The five routing modes, from the docs:

Two subtleties worth knowing: conversations stay with the account that answered them while the vendor's prompt cache is still worth keeping, and intent routing goes further - a small model you choose reads each new turn, so tests can go to the strong model and quick questions to the fast, cheap one. Groups can contain other groups, and the Routing tab shows each decision live.

A central gateway hub with failover paths to provider nodes and a usage meter
When a quota runs out, the next model answers - routing keeps the agent working. Illustration: AI Frontier Post.

6. Watch the meter

Magpie tracks tokens, cache reads and writes, reasoning tokens, and calls - with cost at list price (you can set your own prices per model). Balances and quota windows for every key, plan, and subscription account:

magpie quota   # what's left on every plan

The Usage page shows sessions too - each agent conversation with its cost and title, plus a command to resume it. You can set a limit per key, export usage via OTLP to your own observability stack, and the reset reminder warns you before a window renews with much of it unused.

7. The power moves

export OPENAI_BASE_URL=http://127.0.0.1:3425/v1     OPENAI_API_KEY=magpie
export ANTHROPIC_BASE_URL=http://127.0.0.1:3425     ANTHROPIC_API_KEY=magpie
export GOOGLE_GEMINI_BASE_URL=http://127.0.0.1:3425 GEMINI_API_KEY=magpie

What you built

A single control plane for every coding agent on your machine: any model behind any provider (or subscription) assigned to any agent from one list, routing groups that fail over when quotas run dry, per-session cost tracking with magpie quota, and a local multi-protocol gateway on 127.0.0.1:3425 that any base-URL tool can use.

Honest limitations