SAMPPO: Structure-Aware Mirror Proximal Policy Optimization
Corinna Cortes ⋅ Mehryar Mohri ⋅ Yutao Zhong
Abstract
The alignment of Large Language Models (LLMs) via reinforcement learning from human feedback (RLHF) typically relies on Proximal Policy Optimization (PPO), which uses a uniform clipping hyperparameter $\epsilon$ to maintain a stable trust region. However, we demonstrate that this uniform constraint is ill-suited for the non-uniform semantic topology of natural language. In ``semantically dense'' regions, where candidate tokens act as semantic synonyms within the current context, noisy advantage estimates can destabilize the fixed trust region, leading to *policy collapse*, a severe loss of generative diversity and policy entropy. To resolve this, we propose *Structure-Aware Mirror Proximal Policy Optimization (SAMPPO)*, which introduces an adaptive trust region governed by local semantic density. SAMPPO generalizes the trust-region framework to non-uniform semantic spaces by adapting the mirror-descent geometry to the local action topology. SAMPPO uses a per-timestep *Semantic Isolation Score* to dynamically tighten the clipping range in dense regions and relax it in sparse ones. We provide theoretical foundations for this approach, deriving a *Structure-Aware Monotonic Improvement* bound and establishing a formal link to structure-aware margin-shifted $H$-consistency theory. Empirically, SAMPPO prevents the entropy collapse observed in standard PPO baselines and significantly improves alignment. Evaluated on the Anthropic HH-RLHF and NVIDIA HelpSteer benchmarks across multiple model architectures (Llama-3-8B and Gemma-2-2B), SAMPPO achieves robust, zero-shot generalizable alignment gains. Crucially, it strictly outperforms standard PPO and recent empirical adaptive heuristics (BAPO), demonstrating competitive performance against direct preference baselines including DPO and SimPO on the alignment-diversity Pareto frontier.
Chat is not available.
Successful Page Load