The Model as Its Own Privileged Teacher: A Feedback-Anchored Self-Improvement Loop for Foundation-Model Agents
Abstract
A foundation-model agent that is to keep improving after deployment needs a learning signal that does not depend on demonstrations, on a stronger external teacher, or on a verifier that exists only for narrow domains. We study a loop with three properties: the teacher is the current policy itself, given privileged feedback that the student does not see; the feedback is anchored in expert-authored rubrics rather than in the model's own judgments; and the loop follows a distillation stage with criterion-reward reinforcement learning stages, each producing a small low-rank adapter over fixed base weights. We instantiate the loop on Kimi-K2.7 with 1,465 expert-rubric document-reasoning tasks from ten professional domains and follow three rounds of updates. On the in-domain GDP.pdf holdout, strict pass rises from 11.0% at base to 13.0% after distillation and 24.5% after RL; on 220 out-of-distribution GDPval tasks that require tools and deliverables, zero-credit tasks fall from 54 at base to 29 after RL and 10 after continued RL, mean rubric reward rises from 0.571 to 0.703 to 0.746 over the same checkpoints, tasks improving over base by more than 0.10 rise from 62 to 70 while tasks below base fall from 24 to 18, and messages per task fall from 70.9 to 58.2. Paired trajectories show behaviors maturing with dose: a retry loop with proxy-bypass attempts at one checkpoint becomes a census-first fallback rule at the next; blind scripting becomes unit-tested, idempotent builds with acceptance checks. We describe the harness engineering that makes the privileged self-teacher sound—feedback isolation, exact-token teacher scoring, response-budget reservation, truncation rejection—and report the reliability findings that any such loop must gate on, among them that two evaluation trials of the final checkpoint were discarded because roughly 40% of trajectories ended after a single reasoning-only message. We close with design principles for feedback-anchored self-improvement and with what this study leaves unmeasured, above all forgetting outside the target distribution.