The Illusion of Diversity: Aligning LLM Exploration via Effective Entropy
Abstract
Reinforcement learning with verifiable rewards has shifted the alignment paradigm of large language models from subjective preference to objective correctness, yet a fundamental disconnect persists between exploration diversity and reasoning validity. Existing strategies predominantly rely on ``blind'' entropy metrics, e.g., token-level or global sequence-level entropy, which structurally fail to distinguish between effective reasoning and potential hallucinations. This misalignment creates an illusion of diversity, where high entropy tends to signal chaotic degeneration rather than valuable exploration. To bridge this gap, we introduce \textit{Effective Entropy}, a validity-aware entropy concept that quantifies diversity exclusively within the valid subspace and systematically filters out the noise of incorrect trajectories. Built on this concept, we instantiate VEPO, a simple \textbf{V}alidity-aware \textbf{E}ntropy \textbf{P}olicy \textbf{O}ptimization method that directly incorporates effective entropy into RLVR training. VEPO induces an adaptive attraction dynamic that drives up the probability of under-explored valid trajectories, a mechanism that matches the practical regime of large-scale autoregressive models where sequence probabilities remain extremely small. Extensive experiments on reasoning benchmarks demonstrate that this effective-entropy-driven optimization significantly outperforms GRPO and other entropy-aware variants, showcasing its ability to more reliably cultivate, across diverse tasks, a broad portfolio of valid reasoning trajectories.