Clockwork.io raises $31M to keep AI training alive when GPUs fail — LinkedIn says it saves tens of thousands of GPU-hours a month
Clockwork.io announced $31 million in new funding today, betting that the scarcest resource in AI isn’t compute — it’s the compute you don’t waste. Its fault-tolerance software keeps AI training, inference and reinforcement learning running through GPU, network and server failures, and it now counts LinkedIn and Together AI as production deployments plus an expanded deal with WhiteFiber.
The round, announced today via PR Newswire, was co-led by Premji Invest, Wing Venture Capital and Seligman Ventures, with existing investors NEA and e& Capital participating, and brings Clockwork.io’s total raised to $73 million. Alongside the financing, the company introduced two new capabilities for its TorchPass product — multi-node platform snapshots and fast background application checkpoints — both aimed at preserving distributed jobs without touching training code.
Failure is the normal state
Clockwork’s pitch is blunt, and it borrows its best line from NEA venture partner Greg Papadopoulos: “The one thing that scales perfectly is unreliability: put enough GPUs in one machine and something is always failing.” Large training runs spread across thousands of GPUs that must stay in sync — one dead GPU, one dropped link, one frozen server stalls everything. The release cites Meta’s own report of unexpected interruptions averaging roughly one every three hours during a 54-day Llama 3 training run on 16,384 GPUs, and says conventional checkpoint recovery can take up to 90 minutes, leaving healthy GPUs idle while the job redoes work since its last saved state.
Clockwork sells the alternative as infrastructure, not a guardrail: a software layer between the hardware and the workload. LinkPass reroutes traffic around a failed network link so the job never sees the fault. TorchPass migrates work from a failing GPU to a healthy one so training continues instead of rolling back. CEO Suresh Vasudevan — a veteran of Nimble Storage, Sysdig and NetApp — frames it as a “goodput multiplier”: “Failures are inevitable at AI scale. Losing hours of useful work to them should not be.” (RuntimeWire’s report has more on Vasudevan’s return to startup building after a decade-plus in infrastructure.)
What’s new today: snapshots of a whole job
The two new TorchPass capabilities both capture the state of a running distributed job. Multi-node platform snapshots — billed as an industry first for training — save an entire running job across every node without changes to the training code, so the whole job can be restored after an interruption too large to migrate around. Fast asynchronous application checkpoints are taken in the background while the job runs, so less progress is lost and less compute repeated after a failure.
The second one is aimed squarely at reinforcement learning, which depends on both training and inference: copies of the model generate rollouts, the trainer learns from them, and updated weights must reach the copies before they can generate with the latest version. TorchPass’s fast checkpoints carry updated weights to the rollout replicas sooner, so those replicas spend less time waiting or working from a stale model — while LinkPass keeps them serving through link failures.
The receipts: LinkedIn, Together AI, WhiteFiber
The deployments are the story’s real weight. LinkedIn has deployed LinkPass across its AI infrastructure fleet and, per the release, prevents tens of thousands of GPU-hours of downtime each month. SVP and CTO Infrastructure Raghu Hiremagalur gives the concrete failure math: “Before Clockwork.io, one InfiniBand NIC flap could remove an eight-GPU server from service, while a switch port flap could drain a second server, doubling the impact to 16 GPUs.” Now, he says, disruptive incidents have become “manageable maintenance events.”
Together AI is bringing TorchPass to market as a service on its GPU Clusters; the two companies plan to demonstrate a live multi-node training job at the PyTorch Conference, continuing through injected network and GPU failures without restarting. “Our customers grade us on goodput, the share of their GPU-hours that actually move the model forward,” said Together AI product lead Pavneet Ahluwalia.
WhiteFiber (NASDAQ: WYFI), an existing customer, is expanding Clockwork’s software across its growing global GPU-as-a-service footprint, citing the automated fleet audit: it “validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance,” per CTO Tom Sanfilippo.
Why this round matters — and the caveats
The strongest independent signal comes from Dylan Patel of SemiAnalysis, whose ClusterMAX ratings benchmark GPU cloud providers: “In our ClusterMAX, TorchPass cuts training goodput loss from 14% to under 3% for a gold-rated neocloud.” That’s a third party measuring the thing Clockwork is selling — not a vendor slide. The round itself, covered by Pulse2, confirms the investor roster and the plan to spend the capital on rolling out the fault-tolerance suite across training, inference and reinforcement learning, expanding enterprise adoption, and scaling delivery through cloud partners.
The caveats are the usual ones for a vendor announcement. The GPU-hour savings figures are customer-reported; the release carries no contract values and no per-product deployment breakdown. The “industry first” claim on multi-node snapshots is the company’s own framing, not independently verified. And Clockwork’s roots — founded in 2018 as TickTock Networks, renamed in 2021, relaunched as Clockwork.io in 2022 around co-founder Balaji Prabhakar’s Stanford distributed-systems research — make this a classic infrastructure-company pivot story as much as a funding story.
Still, the direction is hard to argue with: as AI jobs spread across more accelerators, coordination between machines becomes the constraint even when GPUs are available. A software layer that treats failure as the normal state — and prices itself against wasted GPU-hours rather than seats — is selling exactly the problem the largest fleets feel most acutely. More on the technology at clockwork.io.