Adaptive and Neutral Theory of Evolution Strategies for Large Language Models
Soma Yokoi ⋅ Issei Sato
Abstract
Evolution strategies (ES) have recently emerged as a promising alternative to gradient-based reinforcement learning for fine-tuning large language models, but classical zeroth-order optimization theory does not systematically account for the empirical observations driving this success. We formulate the continuous-time limit of Z-score-normalized ES as a unified stochastic differential equation and identify three sequential dynamical regimes characterized by the dominant gradient contribution to the population reward variance; under a basin-local low-rank Hessian assumption, this framework systematically explains (1) the information-geometric optimality of the Z-score heuristic, (2) the dimensionality paradox of sample-efficient ES with populations far smaller than the parameter dimension, (3) the same-task rise-then-decay of the training reward, and (4) catastrophic forgetting of pre-training capabilities. Synthetic experiments and direct verification on Qwen2.5-{0.5B,3B,7B}-Instruct confirm these predictions, matching the parameter drift reported by Abdi et al. (2026) within $\sim 1.5\\%$.
Chat is not available.
Successful Page Load