Scality launches AI Inference Factory: an open-code stack for on-prem AI inference — KV cache from object storage, 14x faster than recompute
Scality announced Thursday the availability of its AI Inference Factory, a supported open-code stack for running AI inference on enterprise-owned infrastructure — serving KV cache from object storage over RDMA at near-GPU-memory latency, 14x faster than recompute in its tests.
Scality, the French-American data infrastructure vendor, announced on Thursday the availability of the Scality AI Inference Factory, a supported open-code software stack for deploying and operating AI inference on infrastructure that organizations own. The pitch is aimed squarely at the pain points of per-token cloud inference: bills that get harder to forecast the more useful a workflow becomes, model versions and quantization that can change at the provider's discretion, and sovereignty concerns about where prompts, documents, and proprietary code get processed.
The target customers are enterprises, government agencies, and neo-cloud providers that want a supported alternative to cloud AI services without assembling, integrating, and maintaining the whole software stack themselves. Scality says the answer for most organizations will be hybrid, with the most critical AI processes running on-premises.
The technical hook: KV cache served from object storage
The centerpiece is Scality's AI Data Infrastructure (ADI) acting as a shared key-value cache that extends KV cache beyond limited GPU high-bandwidth memory. Instead of GPUs recomputing context, they preserve and retrieve it from ADI — a multi-petabyte shared cache that GPUs read at latency Scality says is the same order of magnitude as GPU memory. That matters most for reasoning models and agentic workloads, where KV cache can quickly exceed available GPU memory.
Scality's own test numbers:
- 1.9-second load time for Gemma-3 27B — nearly 10x faster than local NVMe, streamed in parallel across the cluster over RDMA;
- 166 ms warm time-to-first-token on a 14K-token context restore from ADI, 83 ms behind HBM;
- 14x faster KV cache retrieval than recomputation on a 14K-token context, up to 72x on a 439K-token context;
- a KV cache more than 80x larger than a single GPU's memory, keeping up to 1,000 concurrent sessions resumable without recomputation; and
- 97% of network line rate on data transfers between GPUs and storage, with no measurable impact on token generation since context is restored before the first token.
As Scality CTO Giorgio Regni put it, the ADI KV cache is "fast enough to sit in the serving path" — restoring a context from ADI is "the same order of magnitude as GPU memory, and 14 times faster than recomputing it, with the GPUs staying busy the whole time."
Disaggregated serving: a prefill pool and a decode pool
The architecture separates prefill from decode so each can scale independently. One pool of GPUs handles prefill, processing the prompt and writing the KV cache to ADI; a second pool handles decode, reading the cache back and generating tokens. With ADI as the shared cache, any decode GPU can pick up any context, prefill no longer interrupts decode, and both pools run at full load.
Scality cites the academic precedent: DistServe (OSDI 2024) measured up to 7.4x more requests served within the same latency targets from disaggregation, and Moonshot AI's Mooncake reported 75% more requests on Kimi's production traffic from adding a shared KV cache pool on top.
What's in the stack
Alongside the cache, the factory bundles validated open-weight models maintained as part of the supported stack; a disaggregated inference-serving layer; and a control plane that authenticates and meters requests, routes them to the GPUs holding the relevant context, and schedules workloads against SLA targets. A single namespace automatically tiers data across TLC, HDD, and optionally tape.
The stack stays open: every software component ships as open code, and it's compatible with open-source harnesses and agentic frameworks including OpenCode, Hermes, Goose, LangGraph, and Pydantic AI. Validated open-weight models span Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek; the software deploys on standard servers from Dell, HPE, Lenovo, and Supermicro. StorageReview's coverage notes the tools keep working once developers point them at the local endpoint.
IDC's Nataliya Yezhkova called running inference on-premises "a credible option for a growing set of enterprise and public sector workloads, provided the infrastructure can hold and serve model state efficiently at scale" — which is exactly the requirement Scality is addressing across the inference and storage layers.
Scality AI Inference Factory is available now as a software license or as a fully managed service, and the company is demonstrating the solution and its reference architecture at Scality Day in Paris on Thursday.