Dropout-GRPO: Variational Stochasticity for Continuous Latent Reasoning
Wooil Jung
Abstract
Group Relative Policy Optimization (GRPO) relies on diversity among the $K$ rollouts in each group. For latent-reasoning models such as Coconut, deterministic latent recurrence and greedy answer decoding yield identical responses for a fixed prompt and parameters. We investigate whether dropout can supply useful rollout diversity while retaining greedy answer decoding. This setting makes network perturbations the source of exploration, without introducing answer-token sampling. Dropout-GRPO replays the rollout randomness and uses verifier rewards to optimize a group-centered, clipped softmax-score surrogate. On GSM8K, it improves a Coconut baseline from $27.29\\%$ to $29.39\\%$ pass@1, with additional gains on SVAMP and MultiArith. All results use greedy evaluation with dropout disabled. These findings demonstrate that perturbation-based exploration can support verifier-reward post-training of latent-reasoning models while preserving deterministic inference.
Chat is not available.
Successful Page Load