Breadth or Depth? Toward Adaptive Search in Reflective Prompt Optimization
Abstract
Reflective prompt optimizers combine three potential sources of improvement: generating many feedback-guided prompt candidates, carrying feedback information across rounds, and allowing later candidates to inherit earlier edits. To untangle them, we first compare a lineage-enabled optimizer with a budget-matched root-reset policy that generates every revision from the initial prompt. Across the benchmark settings we study, independent revisions often match lineage-enabled search, showing that inheritance is not automatically responsible for the gains of reflective optimization. Independent revisions are often sufficient when feedback is specific to individual examples or can be incorporated in a single step while lineage helps when solving one stage reveals reusable feedback about the next. A controlled experiment suggests that the benefit comes from preserving feedback across rounds: restarting from the original prompt while providing all feedback revealed so far performs about as well as continuing from the latest revision. These findings motivate an adaptive optimizer that allocates compute between parallel root revisions and sequential descendants based on early feedback. We provide a preliminary study toward that goal using small probe with promising signal on our task suite.