The Effects of Minibatches on Automatic Prompt Optimization
Abstract
Prompt optimization methods are increasingly used to improve LLM performance across tasks. Automatic prompt optimizers revise instructions using feedback from labeled examples, often sampled in small minibatches. We investigate one such optimizer, GEPA, to see whether this process makes its prompt revisions overly dependent on the examples used to generate them. In experiments on MATH, newly introduced prompt rules show greater semantic similarity to their update minibatch than child rules that match a parent rule. Showing the reflection model more distinct problems in each update is associated with a smaller gap and a higher probability that an accepted revision also improves GEPA's validation/Pareto score. With an unchanged 100-problem training pool, raising the minibatch size from 10 to 50 increases this probability from 63.5\% to 90.0\%. Training-set subject composition, pool expansion, and repeated copies of the same problems do not yield the same pattern. At the level of individual rules, semantic specificity does not predict leave-one-out utility, and oracle rule removal provides little benefit. These findings identify the composition of a feedback batch as an important design variable in GEPA and suggest testable expectations for feedback-based optimizers such as ProTeGi and TextGrad.