Prime Intellect launches Prime Inference: serverless and reserved GPU serving for open models, tuned for 100 tok/s agents
Prime Intellect has launched Prime Inference, a serving platform for frontier open-source models with serverless endpoints and reserved capacity on its own NVIDIA Blackwell GPUs across multiple datacenters. Announced October 3, 2026, it closes the loop on the company's open training stack — deployed models generate the production traces that feed back into training.

Prime Inference offers two modes: serverless endpoints for variable demand and reserved capacity for sustained workloads, behind one OpenAI-compatible API with automatic failover across datacenters and unified billing with team-level usage tracking. The fleet runs on NVIDIA Blackwell today, with Vera Rubin listed as coming soon. Point any OpenAI SDK at https://api.pinference.ai/api/v1 — the docs carry the full API reference.
This is not a fresh codebase seeing daylight for the first time. Before the public release, Prime Inference processed nearly a trillion tokens per day internally — RL rollouts, synthetic data generation, evaluations, long-running coding agents — and has served large customer deployments in production since January. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22; Prime reports it ranks among the fastest GLM-5.3 endpoints there, with a near-zero tool-call error rate and 100% uptime since launch.

Built for agents#
The target workload is agentic. A typical agent turn adds about 6K new tokens to a 140K-token prompt, reusing most of the conversation history — which makes long-context serving, in Prime's framing, as much a cache-management problem as a compute problem. The company benchmarks that mix with SemiAnalysis's AgentX, which replays multi-turn agent sessions, while also injecting cold arrivals with long prompts. Two numbers are measured: end-to-end tokens per second per user for interactivity, and output tokens per second per GPU for efficiency.
The stack: disaggregation and compression#
The stack combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, developed in partnership with Inferact and NVIDIA, with fixes contributed upstream. Prefill and decode run on separate GPU groups: Dynamo handles routing and orchestration, vLLM runs the model on each group, and once prefill finishes, decoders pull the computed KV cache through NIXL and join the batch. In Prime's tests, separating the two reduced p90 inter-token latency by nearly 40%.
Caching gets the most engineering. Dynamo's KV-aware router picks a prefill worker by weighing cached prefix overlap against queued work; sessions stay on the same decoder between turns to keep KV reuse; Mooncake adds a second cache tier in host DRAM so prefixes evicted from GPU memory can be retrieved instead of recomputed.
The numbers, on GLM-5.3 served on GB200 NVL72 with an interactivity target of 100 tokens per second per user: at that bar, a 1:4 prefill/decode ratio served the most — 66 sessions per prefill group at 101 tok/s per user and 100 output tok/s per GPU. For prefill topology, Prime chose DEP8 — eight data-parallel attention ranks with expert parallelism across the group — which held roughly five times more usable prefix-cache capacity on the same hardware than TEP8. Halving the per-step prefill token budget from 8K to 4K per GPU cut median queue wait from 550 ms to 110 ms and median time-to-first-token by about 20%.

For decode, Prime compressed the 512-value MLA latent with NVFP4 — four bits per value with an FP8 scale per 16 values, while the 64-value positional component stayed FP8. Each cache row shrank from 576 to 352 bytes, lifting per-decoder capacity roughly 50%, from 1.09 million to 1.63 million cached tokens at the same memory budget. A native sparse-MLA decode kernel unpacks the compressed rows on-chip as attention needs them, keeping NVFP4 as the storage format while computing in FP16 with FP32 accumulation.
Reliable tool calls#
Agents fail when tool calls carry wrong names or broken arguments, so Prime treated reliability as a first-class target. The team contributed a structural-tag builder to Dynamo for GLM's tool format, which vLLM then uses with xgrammar to mask tokens that would violate the tool schema — and fixed parsing bugs along the way, including < being decoded into < inside code.
Why it matters#
Prime enters a crowded inference market — Together AI, Fireworks, and Baseten all serve the same open models, and all already offer GLM-5.3. Prime's differentiation pitch has two parts: an open, disclosed serving stack rather than a black box, and the training-loop connection. Prime already ships post-training tools — prime-rl, verifiers, sandboxes, RL environments — and serving is the missing stage: deployed models generate production traces that can feed back into training, the continual-learning loop the company's whole stack is built around.
One open question: per-model pricing is not yet published in the docs, so the economics of the service can't be compared against rivals yet. What is published is the engineering — the launch post walks through every optimization in depth — which is itself the pitch: this is inference as an open infrastructure project, not a managed endpoint you are meant to treat as magic.