Beyond Clipping: Signed Logarithmic Smoothing for Policy Optimization
Jihun Yun ⋅ Sungjoon Yoon ⋅ Beomhan Baek ⋅ Minhak Song ⋅ Jongha (Jon) Ryu ⋅ Kwang-Sung Jun
Abstract
Policy optimization algorithms that reuse rollouts build updates of the form $\rho A$, where $\rho$ is the importance ratio from the current policy to the rollout policy and $A$ is an advantage estimate. Recent RL algorithms such as PPO and GRPO rely on clipping $\rho$ for stable policy updates, which introduces further nonsmoothness in the optimization landscape. \emph{Is clipping the best we can do?} In this paper, we propose signed logarithmic smoothing (Signed LS), $\psi_\beta(z) = \mathrm{sgn}(z)\,\beta^{-1}\log(1+\beta|z|)$, applied to the full weighted update rather than to the ratio alone. Building on this surrogate, we introduce KRAFT, a GRPO-style policy-optimization objective that replaces the clipped surrogate with Signed LS. For theoretical justification, we consider an adaptive off-policy contextual-bandit setting: we compare the deviation structure of Signed LS with clipping-style surrogates, prove surrogate fidelity for Signed LS, and obtain regret transfer against a fixed comparator policy. The comparison highlights a structural distinction: ratio clipping introduces explicit tail-bias and threshold-dependent concentration terms, whereas Signed LS controls distortion through a smooth second-order quantity. Empirically, our contextual-bandit experiments provide evidence of these effects. In RLVR experiments on math-reasoning benchmarks and multiple LLM backbones, we observe that KRAFT consistently shows lower KL from the reference policy and higher entropy throughout training without incurring accuracy loss compared to GRPO.
Chat is not available.
Successful Page Load