Fault Tolerance in Transformers Favors Output Alignment over Hidden-State Consistency
Haonan Tan ⋅ Gene Wen ⋅ Yuxing Han
Abstract
Can a large language model be \emph{trained} to tolerate silent hardware faults, and if so, \emph{where} in the networkshould the consistency signal be enforced? We answer both questions and find that the second one matters more than the first. Injecting tile-level GEMM faults during training and adding a KL consistency term on the output distribution reduces fault-induced perplexity degradation from $12.9\%$ to $0.21\%$ on GPT-2 Small---a $60\times$ reduction (two-seed mean). The placement of the consistency target is decisive: a direct hidden-only objective at the fault site can worsen degradation to $32.7\%$, and fair bounded hidden controls with corrupted CE improve substantially but still trail output KL in an 80k-step screen ($1.74$--$1.76\%$ vs.\ $0.84\%$). A conditional constraint hierarchy and gradient-alignment probes explain why output-level consistency is more permissive and easier to optimize than direct early hidden-state consistency, without requiring that every hidden-alignment formulation fail. The output-over-hidden advantage persists at 8B scale, the GPT-2 ordering reproduces on held-out corpora, and generic robustness methods---SAM, R-Drop, and SAF---fail to close the gap. Pairing the trained model with simulated algorithm-based fault detection yields a bit-flip-to-erasure interface with a $1.10\times$ perplexity ratio, where unprotected inference produces NaN.
Chat is not available.
Successful Page Load