Degeneration Detection Is Not Early Warning: False-Alarm Evaluation with ONWARD
Abstract
Internal language model (LM) representations are known to carry signals of repetitive degeneration and, in some settings, to anticipate its onset. We study whether token-level linear probes can turn these signals into alarms that fire while a response is still being generated. For this purpose, we train probes on residual-stream activations and evaluate their alarms relative to the annotated onset of degeneration under controlled response-level false-alarm rates. Across two instruction-tuned LMs, strong probes are nearly indistinguishable by AUC and accuracy, across training strategies and layers alike, despite substantial differences in early-warning behavior. For this purpose, we introduce ONWARD, an evaluation protocol that assesses probe alarms relative to the annotated onset of degeneration under controlled response-level false-alarm rates. ONWARD expose these differences and give a more informative basis for probe selection. At a 5% false-alarm budget, the selected probes achieve strong post-onset coverage and, among alerted degenerate responses, fire before the annotated onset in a majority of cases. Low-rank adaptation raises pre-onset warning coverage without raising the early-warning rate, suggesting it makes an existing signal more consistently linearly readable.