When Rank Rises as LLMs Degrade: Stress-Testing Spectral Monitors for Continual Post-Training
Zhaohui Wang
Abstract
Post-training adapts a language model in a non-stationary environment, and practitioners increasingly monitor representation health — RankMe and related spectral statistics — to decide when adaptation has gone wrong. Such monitors carry an implicit directional assumption inherited from the self-supervised vision literature: rank goes down when representations degrade. We show that this assumption is not safe under LLM post-training. In a controlled sweep over degradation modes (Qwen3-0.6B, four modes × three seeds), a data-duplication regime that degrades held-out loss by 75% relative to a healthy run drives RankMe above healthy — in its original uncentred form as well as the centred variant (13.5 pooled s.d.) — and the covariance effective rank to nearly twice healthy: damage there is spectral dispersion, not collapse, so a one-sided monitor scores the worst checkpoint as the healthiest. A learning-rate misconfiguration, by contrast, moves RankMec and k95 in the conventional direction while the original uncentred RankMe is inconsistent across seeds — so the sign is a property of the regime–statistic pair, and cannot be fixed by recalibration. We also separate two statistics the literature often conflates: RankMe normalises singular values whereas covariance effective rank normalises eigenvalues, and on raw intermediate-layer hidden states the massive-activation phenomenon pins the latter near 1 out of $d$ on an unmodified pretrained checkpoint while leaving the former with usable range. We then stress-test the natural repair — two-sided, multi-channel sequential monitoring with a multiplicity-aware false-alarm target — with calibration and test data held strictly apart. In a pre-registered shared-prefix, leave-one-seed-out evaluation on Qwen3-0.6B, the two-sided ensemble detects all three damage regimes in every fold within 10–60 steps of the fork, and the firing direction separates spectral dispersion from downward-rank damage; but it never precedes the held-out probe loss, and with two calibration seeds it does not achieve zero false alarms on the held-out healthy seed. We report this negative result in full: spectral monitoring diagnoses the failure regime, it does not warn earlier than a held-out loss, and validity claims made without a held-out healthy seed should not be trusted.
Chat is not available.
Successful Page Load