Internal Stability Is Not Evidence of Replication: Evaluating Attention-Based Self-Correction Signals
Abstract
Internal stability should not be interpreted as evidence of replication: an evaluation pipeline can appear reliable within an exploratory sample while failing under additional data or stricter validation. We test whether attention activations immediately before a language-model reasoning pivot distinguish successful corrections from non-corrective pivots. In Qwen2.5-1.5B-Instruct, an initial 75-record sample yielded out-of-fold ROC-AUC 0.720 (average precision 0.411), but after adding 52 later-collected records the estimate fell to 0.579 (average precision 0.218). Expanded prompt-grouped cross-validation yielded ROC-AUC 0.473 and average precision 0.174, essentially its 0.173 prevalence baseline; training on the original 75 and evaluating only on the later 52 yielded ROC-AUC 0.392. Across 30 prompt grouped split seeds, the initial estimate was internally stable (mean ROC-AUC 0.702), whereas the expanded estimate was near chance (mean 0.516). No head survived FDR correction. A pre-pivot TF–IDF baseline matched or exceeded the initial attention score but likewise failed after expansion and on later records, showing that the early predictability was not unique to internal activations. Reconstructed prompt-plus-generation context and a secondary Gemma-2-2B-IT check did not rescue the conclusion. These results do not show that internal correction signals are absent; they show that this high-dimensional protocol, at the available sample sizes, does not support a stable, transferable evaluator.