nnTrace: Detecting and Localizing Silent Bugs in Distributed Training
Abstract
Distributed training is essential for scaling LLM training across thousands of GPUs. However, as distributed training requires complex implementations, they are prone to silent bugs, which do not produce explicit error signals but lead to in correct training outcomes. Common debugging practices based on monitoring training loss or gradient norm curves are slow or unable to detect bugs and also do not help localize bugs. We design and implement nnTrace, the first systematic differential testing system for detecting and localizing silent bugs in distributed training. nnTrace aligns intermediate tensors from distributed training with those from a trusted reference implementation. To properly compare the floating-point values in the corresponding tensors, we propose a novel mathematical analysis that provides a guideline for setting tolerances, enabling nnTrace to distinguish bug-induced errors from numerical errors. Experimental results demonstrate that nnTrace effectively detects 11 existing bugs and 3 new bugs in the widely used Megatron-LM framework. nnTrace is effective in various training recipes, including low-precision recipes involving BF16 and FP8. Notably, Megatron-LM has already adopted the method proposed by nnTrace in its development workflow. Our code is available at https://anonymous.4open.science/r/NECK-3C61/.