Does Search Strategy Matter? A Study of Orchestration in LLM GPU Kernel Optimization
Abstract
An LLM agent can rewrite a GPU kernel, but a complete system must also search: pick a candidate, invoke the agent, verify the result, and decide what to try next. We call the LLM that proposes each edit the worker and the loop that chooses what it edits next the orchestrator. Prior kernel-agent systems vary both at once, so the value of the search strategy alone is unknown. We isolate it, holding the worker, verifier, tasks, and scoring fixed while varying only the orchestrator's strategy across six schedulers, ~1,500 trials whose every edit is compiled, checked for correctness, and timed on real hardware. The biggest lever is not the search strategy: which kernel you pick to optimize explains 83-95% of the variation in speedup, the strategy at most 4.5%. Within that margin the aggregate ordering is consistent. Starting fresh from the original kernel each time is the weakest use of a fixed budget in every setting we tried; depth-first refinement so far is the strongest. Per kernel, though, no strategy wins outright.