Geometric Prompt-Trajectory Planning for Test-Time Scaling
Zhengqi Pei ⋅ Anran Zhang ⋅ Qingming Huang ⋅ Shuhui Wang
Abstract
Test-time scaling (TTS) improves large language model (LLM) reasoning by spending additional inference-time compute, most commonly through width-heavy repeated sampling and aggregation. However, this strategy scales cost almost linearly with the number of calls and can remain brittle when the sampled trajectories share the same failure mode. We introduce Maze, a test-time scaling framework that shifts computation away from repeated target-model calls and toward LLM-free planning over ordered few-shot exemplars. The central idea is to treat an ordered exemplar sequence as a prompt trajectory embedded in a latent geometric space. Instead of asking the LLM for many independent attempts, Maze first ranks candidate prompt trajectories using cached encoder representations and a path scorer, ${\it i.e.}$, a lightweight learnable metric, then spends one call (or a few calls) only on the most promising trajectories. This design preserves black-box deployability, since the planner is decoupled from the target LLM and can be implemented with a separate frozen encoder. Across various challenging reasoning benchmarks, single-call Maze consistently outperforms strong single-call prompt-optimization baselines, and MultiMaze recovers a large fraction of width-heavy TTS gains with substantially fewer computational cost measured by number of LLM calls and latency. We further present a theoretical perspective showing when prompt-path ranking can recover much of the benefit of best-of-$N$ style scaling: if high-utility prompt trajectories are sufficiently rankable and sufficiently covered by the explored pool, a deeper planner-side search can substitute for a substantial amount of width-heavy sampling.
Chat is not available.
Successful Page Load