Learning Where to Constrain: Meta-Learned Trust Allocation for Mathematical Reasoning
Abstract
Long mathematical solutions expose a weakness in uniform policy-update constraints. Early token changes affect a long suffix, while late changes affect fewer future states. Prefix divergence-based methods like CPPO encode this asymmetry with a fixed position schedule and prefix budget. We introduce Meta-TA, a bilevel method for learning both choices from post-update performance on disjoint prompts. A direct one-step design fails when behavior and current policies match because the constraint starts inactive. We prove the resulting hypergradient equals zero. A detached warm-up creates controlled policy drift. A soft-gated step then supplies the meta-gradient. On Qwen3-8B-Base, five frozen-retraining seeds raise mean AIME24/25/26 Avg@16 from 29.01 for CPPO to 31.95 for Meta-TA with closely matched realized KL. At a 16k response cap, Meta-TA lowers normalized reasoning effective horizon from 0.475 to 0.430 and importance-ratio variance from 0.0534 to 0.0475. These results link adaptive trust allocation to higher accuracy and more localized updates in long-horizon mathematical reasoning.