Self-Sealing Reward Models: When Feedback Cannot Refute Itself
Manoj Saravanan ⋅ Rohit Kumar Salla
Abstract
Self-rewarding post-training changes both the policy that generates future comparisons and the evaluator that scores them, so learning changes the experiment used to validate its own reward. We characterize when this loop becomes self-sealing. For comparison coverage $C_t$ and evaluator sensitivity $S_t$, define $G_t=S_t^\transpose C_tS_t$ and $\Ieff_T(h)=\sum_{t\leq T}h^\transpose G_th$. We prove that, for any target displacement $h$, the event $\{\Ieff_\infty(h)<\infty\}$ is exactly the equivalent component of the two infinite transcript laws, while its complement is singular. In deterministic designs, target-policy identifiability is equivalent to the affine finite-information fiber remaining inside one normal-fan cell; any positive equivalent component yields a nonzero minimax loss floor and linear deployment regret. For explicit KL-proximal policy and coverage-weighted evaluator updates, we derive the sharp boundary $\alpha_\pi+2\kappa_E=1$: above it, support, sensitivity, and one-round information stay positive, but total information is finite and policy-incompatible worlds are equivalent. Finally, global repair requires recurrent trusted information of minimum rank $\dseal=\dim(\Eseal/(\Eseal\cap\Dpol^\perp))$, with time-varying repair characterized by persistent excitation on the sealed policy quotient.
Chat is not available.
Successful Page Load