Label-Free Consistency Correction for Weak-to-Strong Generalization
Abstract
Most explanations of weak-to-strong (W2S) generalization focus on why a strong student may fail to exactly imitate weak-label errors. We study a complementary mechanism: before student training, weak-teacher predictions may be inconsistent along semantic constraints specified independently of the weak model. When such constraints are valid for the target, these inconsistencies give label-free lower bounds on the risk reduction achievable by correcting weak labels, rather than merely indicating unstructured prediction error. We formalize this mechanism using Markov and Laplacian operators. Under exact target validity, Markov trajectory inconsistency equals the squared-risk gain from smoothing; under approximate validity, it remains a conservative lower bound after an explicit validity penalty. We then solve the associated minimax correction problem, obtaining a resolvent correction that continuously attenuates constraint-violating directions and reduces to the identity when validity uncertainty is too large. These results yield validity-adjusted correction (VAC), a split-sample pipeline that selects corrections from weak-teacher queries and predeclared validity budgets before training the strong student. Controlled graph-signal experiments and a CIFAR-10 cats-versus-dogs W2S experiment are consistent with the proposed mechanism: validity-adjusted inconsistency predicts useful correction and transfers to student gain, whereas high raw inconsistency under invalid controls does not provide such evidence.