Spreading AI training across datacenters creates new networking hurdles
Large AI models are increasingly trained across multiple datacenters instead of a single facility. Power and space limits are pushing companies like Google, Microsoft, AWS, and Meta to link clusters over long distances. Training still requires frequent synchronization and the exchange of numerical data such as gradients among thousands of accelerators.
The move reflects physical limits. One site or campus may lack enough power and space for modern models. Cisco estimates today’s training clusters can involve tens of thousands of GPUs; Epoch AI researchers project that by 2030 the largest frontier runs could draw 4–16 GW. Distributing work lets operators use regions with more power, room, and fewer planning constraints.
Training adjusts billions of weights through repeated prediction, error measurement, and updates. This work is split across thousands of accelerators—GPUs, AWS Trainium, Google TPUs—that must periodically exchange gradients and intermediate results. Synchronous jobs cannot let faster groups continue while slower ones lag; the collective exchange must finish before computation resumes, so network delay can become training delay.
This shift may concentrate frontier AI development among cloud providers and well-funded labs able to secure long-distance network capacity, power, and land. Communities hosting expanded datacenters could see new infrastructure investment and electricity demand, while researchers and smaller firms might face higher barriers to training comparable models. Everyday users may eventually feel these changes through the cost, availability, and capabilities of AI services, though the scale and direction remain uncertain.