Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
Abstract
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is dominated by the GRPO and DPO families, yet these methods often update too aggressively on the high-variance samples from agentic tasks. We propose Follow the Winners (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, obtaining a more conservative policy-learning method. Unlike prior exponentially concentrating approaches, FTW induces polynomial concentration in the order statistic of returns. Through a control-as-inference lens, we show that FTW induces a mild risk-seeking bias that scales linearly with return uncertainty: less aggressive than DPO’s quadratic risk bias, yet more exploratory than GRPO’s risk-neutral profile. In deterministic-reward agentic settings, where uncertainty is primarily epistemic, this linear risk bonus encourages knowledge-seeking without entrenching on noisy samples. Consistent with this perspective, FTW matches GRPO on the perfect-information Sokoban domain while outperforming it on the imperfect-information Search-R1 domain.