Clockwork.io Secures $31M to Prevent Costly AI Training Interruptions from Hardware Failures

Clockwork.io announced a $31 million funding round to expand its infrastructure software that enables AI workloads to continue operating when GPUs, networks, or servers fail. The company's technology, deployed by LinkedIn and other major platforms, prevents tens of thousands of GPU-hours of downtime monthly by rerouting traffic and shifting work away from failing components. With distributed AI training jobs now stretching across thousands of GPUs, hardware failures can cause expensive delays and force systems to restart from earlier checkpoints.
Clockwork.io's infrastructure software addresses a critical pain point in modern AI development: the economics of large-scale distributed training. When training jobs span thousands of GPUs across multiple servers and networks, even brief hardware failures cascade into significant losses—idle equipment, repeated computational work, and extended training timelines. The company's dual-product approach handles different failure types: LinkPass manages network disruptions by redirecting traffic, while TorchPass allows GPU workloads to migrate away from failing hardware without forcing full system restarts.
Real-world deployment validates the technology's value. LinkedIn's adoption prevents tens of thousands of GPU-hours of monthly downtime, translating directly to cost savings and faster model development cycles. The funding round's backing from established venture firms and adoption by major platforms like Together AI and WhiteFiber suggests confidence that fault tolerance will become standard infrastructure in AI development.
Clockwork.io's technology could reshape AI infrastructure economics by reducing the operational overhead of training large models. As AI development scales, fault tolerance may shift from optional optimization to essential capability, potentially democratizing access to large-scale training by making it less wasteful for organizations with constrained resources. However, widespread adoption would primarily benefit companies and institutions already operating at scale, potentially reinforcing competitive advantages for well-funded AI developers.