The Holonomy of Thought: Taming Compounding Error via Curvature Regularization
Abstract
Foundation-model agents increasingly improve through repeated verifier-guided updates on mathematical reasoning, code generation, and tool-augmented web tasks. Yet endpoint rewards constrain final outcomes rather than the internal computation along sampled trajectories. Local deviations can therefore accumulate over long executions and destabilize policy optimization. We formalize this path dependence as discrete holonomy in hidden-state transport. We prove a local-to-finite control chain: local transport residuals telescope into full-path Jacobian disagreement; transported curvature accumulates into the relative holonomy characterized by this disagreement; and the disagreement controls finite endpoint error. This connection motivates Path-Jacobian RLVR, which directly regularizes transport disagreement between same-prompt rollouts using efficient shared-probe vector-Jacobian products. Experiments across mathematical reasoning, code generation, and interactive decision-making settings show that Path-Jacobian RLVR generally outperforms the PPO baseline across multiple models and tasks. Ablations over Jacobian estimator budgets and model architectures further demonstrate improved training stability and the generality of the gains across both discrete and continuous trajectory models.