
Every coding agent keeps its model in its own file, in its own format, with its own keys and base URLs. Your Claude Code subscription works in Claude Code and nowhere else. Your cheap DeepSeek key works in the agent that lets you paste a base URL, and nowhere else. When a quota runs out at 3 pm, you start editing config files - and every agent has a different idea of which vendors it even supports.
Magpie puts all of it in one place. It is a local model gateway - one process, one screen - that runs Claude Code on Kimi, Codex on DeepSeek, Gemini CLI on GLM, or OpenCode on your ChatGPT plan. It ships a gateway on 127.0.0.1:3425 that speaks the OpenAI Chat, OpenAI Responses, Anthropic Messages, and Gemini protocols, translating between them with streaming, tool calls, and reasoning intact. Pick any model for any agent from a single list; magpie rewrites only that one key in the agent's own config file, atomically, leaving your comments and formatting alone. It is MIT-licensed, written in Go, runs on macOS, Windows, and Linux, and picked up 5,100+ stars in about 12 days.
What you get at the end: every agent on your machine pointed at whatever model you want, plus routing groups that fail over to the next model when a quota runs dry - so your agent never sees the error. Here is the whole thing, hands-on.
The one-liner from the README:
curl -fsSL https://usemagpie.ai/install.sh | sh
Alternatives: download the app from usemagpie.ai (Mac builds are signed and notarised, and every build updates itself), go install github.com/yetone/magpie@latest, or the Docker image ghcr.io/yetone/magpie. Behind a firewall, the installer takes --proxy or --mirror.
Open magpie and go to Providers -> Add provider, pick a preset, paste a key - done. Or do it from the terminal:
magpie provider add deepseek sk-...
Presets include Anthropic, OpenAI, Google Gemini, DeepSeek, Kimi, Zhipu GLM, MiniMax, StepFun, Qwen, Mistral, Groq, xAI, OpenRouter, Together, Fireworks, Ollama, LM Studio - plus any OpenAI-compatible or Anthropic-compatible URL. The model list comes from the vendor itself, with names and reasoning levels filled in from models.dev, so a model released this morning shows up on the next refresh. No model list is baked into magpie. Already configured elsewhere? Import reads the providers you set up in CC Switch, Claude Code, Codex, and Alma.
On the Agents page, click a model value and a filtered list opens holding every model of every provider as provider/model. Pick one - magpie rewrites just that key in the agent's own config file, atomically. The CLI form:
magpie claude deepseek/deepseek-v4-pro # Claude Code on DeepSeek
magpie codex moonshot/kimi-k2.5 # Codex on Kimi K2.5
Start a new agent session and it uses the new model. The same list and the same click work across all 35+ supported agents - Claude Code, Claude Desktop, Codex, Gemini CLI, OpenCode, Cursor CLI, Cline, Zed, VS Code Chat, Copilot CLI, and the rest.

This is the part that surprised me. An agent you are already signed in to becomes a provider in magpie - your Claude, ChatGPT, Copilot, Gemini (Code Assist), or Grok sign-in works in every other agent, with no key to copy. Several accounts per subscription are supported, with failover between them:
claude binary and bridges your agent's tools over MCP.pi provider packages from npm run in magpie as they do in their own apps - magpie plugin add opencode-gemini-auth, then magpie plugin login google-plugin for its own sign-in flow.A routing group is several models that an agent picks as one. When one hits a rate limit or runs out of quota, the next one answers - your agent never sees the error:
magpie group add "Opus anywhere" models=claude/claude-opus-5-5,copilot/claude-opus-5.5 routing=smart
magpie claude group/opus-anywhere # fails over between subscriptions
The five routing modes, from the docs:
smart - of the subscriptions with quota left, use the one whose allowance renews soonest, so less is lost at the reset.order - use the first model until it cannot answer, then the next.rotate - move to the next member on each turn.usage - use the least-used member first.pace - use the account with the most of its week left per hour until its reset.Two subtleties worth knowing: conversations stay with the account that answered them while the vendor's prompt cache is still worth keeping, and intent routing goes further - a small model you choose reads each new turn, so tests can go to the strong model and quick questions to the fast, cheap one. Groups can contain other groups, and the Routing tab shows each decision live.

Magpie tracks tokens, cache reads and writes, reasoning tokens, and calls - with cost at list price (you can set your own prices per model). Balances and quota windows for every key, plan, and subscription account:
magpie quota # what's left on every plan
The Usage page shows sessions too - each agent conversation with its cost and title, plus a command to resume it. You can set a limit per key, export usage via OTLP to your own observability stack, and the reset reminder warns you before a window renews with much of it unused.
magpie save work && magpie use work - save every agent's setup under one name ("Budget", "Focus") and switch them all at once.magpie tui (the whole thing in a terminal), magpie web, and a plain CLI.export OPENAI_BASE_URL=http://127.0.0.1:3425/v1 OPENAI_API_KEY=magpie
export ANTHROPIC_BASE_URL=http://127.0.0.1:3425 ANTHROPIC_API_KEY=magpie
export GOOGLE_GEMINI_BASE_URL=http://127.0.0.1:3425 GEMINI_API_KEY=magpie
A single control plane for every coding agent on your machine: any model behind any provider (or subscription) assigned to any agent from one list, routing groups that fail over when quotas run dry, per-session cost tracking with magpie quota, and a local multi-protocol gateway on 127.0.0.1:3425 that any base-URL tool can use.
claude binary and plugins run their own sign-in flows - if a vendor changes its login, that provider can break until magpie adapts.