A Foundational Model System for Datacenter Machine Repairs
Abstract
The rapid scale-up of AI infrastructure has made datacenter reliability a cornerstone of the cloud business economy. Hardware failures in hyper-scale AI clusters do not merely disrupt individual machines; they stall multi-million dollar training workloads, and so the downtime of individual machines can cause outsized economic losses. Minimizing machine downtime, therefore, is more important than ever. To address this, we present DrWatson, an AI-driven hardware diagnosis system that pioneers a foundation-model approach to automated datacenter maintenance. Deployed at production scale within a major AI hyperscaler, DrWatson yields substantial efficiency gains across diverse compute architectures. Real-world evaluations demonstrate that it reduces production downtime by 18.3% on GPU platforms and 16.1% on custom AI accelerators. Furthermore, it significantly accelerates hardware deployment, cutting quality assurance (QA) time by 28.3% for GPUs and 25.1% for custom accelerators. These results establish DrWatson as a field-tested solution for maximizing the reliability and cost efficiency of next generation AI fleets.