Beyond Correctness: Robustness-Driven Evolutionary Self-Training for Large Language Models
Abstract
Outcome-based post-training methods for mathematical reasoning rely on binary feedback, treating all correct trajectories equally. However, this masks a critical distinction: some correct paths are structurally brittle and prone to cascading errors under inevitable autoregressive sampling noise, while others are robust and maintain their correctness despite these natural decoding variations. Optimizing indiscriminately over these brittle solutions leads to inefficient learning, as the model wastes capacity memorizing brittle reasoning strategies. To address this, we propose Trajectory Robustness-Driven Evolutionary Self-Training (TREST). Instead of relying on binary feedback, our framework explicitly prioritizes trajectory robustness by integrating evolutionary algorithms with supervised fine-tuning. Crucially, our analysis reveals that trajectory robustness serves as a strong natural proxy for genuine mathematical insight. By conceptualizing reasoning paths as evolving individuals within a population, this evolutionary process naturally marginalizes brittle, brute-force calculations in favor of robust, insight-driven strategies. By effectively internalizing these strategies, TREST achieves superior reasoning performance. Experiments on complex mathematical benchmarks demonstrate that under aligned budgets, TREST outperforms outcome-based baselines, establishing a highly effective alternative to conventional post-training paradigms.