VDURA V12 is live: hyperscaler-style storage for AI factories — 8 to 100,000 GPUs at 2x+ performance per watt
VDURA has taken Data Platform V12 generally available: multi-tenant, API-driven storage qualified on Supermicro that scales from 8 to 100,000 GPUs on one software stack — with persistent KV caching for inference, a claimed 2x+ performance per watt, and a pitch that storage, not GPUs, is where neoclouds win their margins.

VDURA is going after the least glamorous bottleneck in the AI factory business: storage. The company announced today the general availability of the VDURA Data Platform V12, a release it describes as turning its platform into a multi-tenant, API-driven storage service for GPU clouds and AI factories. V12 ships as a qualified solution on Supermicro Building Block Solutions and scales from 8 to 100,000 GPUs on a single software stack.
All the performance and economics claims below come from VDURA's own announcement, released today via Business Wire. They have not been independently verified — but the specs are unusually specific, which makes this one worth reading closely.
The margin layer of the AI factory#
VDURA's pitch is aimed squarely at neoclouds — the GPU-cloud operators that rent out accelerators by the hour. The argument: storage determines how much of that GPU capacity actually turns into revenue. "Every accelerator waiting on a checkpoint, a model load, or a cold read is margin lost," the company says. "Every kilowatt spent on storage is a kilowatt not spent on compute."
CEO Ken Claffey's quote in the release reads like a requirements list from those operators: "keep the GPUs fed, isolate the tenants, automate everything, expand capacity for cold data without a second system, and move that data back to flash the moment it warms up for extended context." His punchline: "V12 is that list, shipped. It is the same mixed-fleet, software-defined model the hyperscalers run inside their own clouds, delivered on Supermicro systems our customers already buy. Every watt and every rack unit we give back is another GPU the operator can put into service."
What's new in V12#
V12 extends VDURA's HYDRA architecture — its High-performance, Yield-optimized, Distributed, Resilient Architecture — with the capabilities multi-tenant GPU infrastructure requires:
- Multi-tenant by design. Per-tenant quality of service, namespaces, encryption keys and VLAN isolation on one shared fleet, so a provider can carve a single storage pool into hard-walled tenant services with capacity and performance guarantees.
- API-first automation. REST APIs, Kubernetes CSI and infrastructure-as-code tenant provisioning, so storage is deployed, provisioned and billed through the same pipelines as the rest of the GPU cloud.
- Context-Aware Tiering™. Data lands on the right media automatically as access patterns shift between training and inference: roughly 90% of files stay on flash while roughly 90% of capacity settles on HDD, in one platform with no stub files, no rehydration steps and no manual tuning.
- Persistent context for inference. A KV cache that outlives the pod. Sessions resume instead of prefilling again, delivering faster first tokens and more concurrent users on the same GPUs, at flash cost rather than recompute cost.
- RDMA data paths. Direct GPU-to-storage transfers with the CPU out of the path. The DirectFlow™ parallel client takes roughly 191 MB of DRAM and zero cores from the GPU node.
- Elastic Metadata Engine. VeLO™ metadata acceleration of up to 20x improvement, 225,000 creates and deletes per second per Director, and billions of metadata operations per second in aggregate.
- File and S3 in one platform. An S3 object is a file in the volume, not a copy of one. No staging copies between ingest, training, inference and archive.
- Snapshots and SMR HDD optimization. Instantaneous, space-efficient snapshots for checkpoints and operational recovery, and SMR HDD unlocking 25 to 30% more capacity per rack.
- End-to-end encryption. AES-256 at rest and in flight, with KMIP key management per tenant.
- Self-healing resiliency and VDURA Sentinel™. Failure domains as small as a single VPOD™, no manual rebuilds and no downtime windows, backed by VDURA Sentinel™ proactive support that opens the service request, with the diagnosis attached, before the customer sees a fault.

Qualified on Supermicro#
The qualified configuration is all Supermicro, and specific: the AS-1116CS-TN — a 1U system with a single AMD EPYC 9005 series processor and 12 NVMe bays — serves as both the VeLO Director node and the all-flash F-Node, alongside the AS-2015HS-TNR hybrid storage node and the CSE-947HE2C 4U 90-bay JBOD for the mixed-fleet data plane. Every node connects with RDMA straight to the GPU nodes, with no dedicated back-end storage fabric.
The economics framing: clusters grow online from three nodes to thousands, and flash share is a dial rather than a fork. A single platform at roughly 20 PB usable spans a 35x performance range — from a capacity-optimized mixed fleet at 2% flash and 18 kW to an all-flash configuration at 1,000 MB/s per TB. VDURA claims 2x+ performance per watt and more than 60% lower total cost of ownership than competitive architectures at the same feed rate.
The track record#
VDURA is leaning on lineage here: 25 years of parallel file system engineering, from PanFS to VDURA, with more than 1,000 production deployments in over 50 countries and namespaces of more than 1,500 nodes. V12 was named AI Data Management Solution of the Year in the 2026 AI Breakthrough Awards. The company will demo V12 on Supermicro at Ai Everything Abu Dhabi, October 6–7 at the ADNEC Centre (Booth H3-D45) — with partner Hafþór "Thor" Björnsson, yes, that Thor, alongside the team.

Availability#
V12 is generally available today for all V5000 class systems and as an upgrade for V11 customers. VDURA Sentinel™ is included; proactive service requests and parts dispatch are delivered through VDURACare Premier, a 10-year support offering covering hardware, software and 24x7 expert response under a single contract.
Why this matters#
The neocloud boom has made GPU supply the headline, but the infrastructure that keeps GPUs fed — storage, networking, power — is quietly where operators differentiate. Checkpointing alone can stall a training run for hours if the storage tier can't absorb and restore model state fast enough, and inference economics increasingly turn on KV-cache management, not raw FLOPs.
A storage platform that makes multi-tenancy, tiering and persistent KV cache first-class features — on commodity Supermicro hardware the operators already buy — is a plausible answer to the neocloud margin question. The caveat that travels with every vendor release still applies: these are VDURA's numbers, independently unverified for now. But with AI factories multiplying and watts as scarce as GPUs, the bet that storage decides who profits is worth watching.