Stable and Granular Policy Optimization for Generative Recommendation
Abstract
Generative recommendation (GR), which directly generates sequential semantic IDs for items, has recently attracted increasing attention in recommender systems. Given that core recommendation metrics are inherently evaluated at the sequence level, reinforcement learning (RL) serves as a natural optimization framework. However, adapting recent group-based RL algorithms to GR presents notable challenges: token-level methods such as GRPO exhibit much more severe training instability, while sequence-level methods such as GSPO enforce uniform token weighting and thus ignore the heterogeneous roles of tokens. To address this dilemma, we propose Stable and Granular Policy Optimization (SGPO), a group-based RL algorithm tailored for GR. SGPO maintains a sequence-level ratio to promote optimization stability while injecting fine-grained token-wise advantages. By incorporating token log-probabilities and reward signs, SGPO explicitly upweights high-confidence tokens in positive trajectories to reinforce successful retrieval, while strictly penalizing them in negative trajectories to encourage exploration, and uses an adaptive scaling factor to bound the variance of reshaped advantages. Experiments on two public benchmarks and a billion-scale industrial dataset show that SGPO consistently outperforms six strong RL baselines and, when deployed on a commercial platform, achieves a 2.36% lift in click-through rate (CTR) and a 1.95% increase in advertising revenue in online A/B testing.