Evaluating Depth and Breadth in Test-Time Scaling for Compositional Visual Generation
Abstract
Test-time scaling is increasingly used to improve visual generation, yet it remains unclear how different scaling strategies should be evaluated when they trade off accuracy, latency, throughput, compute cost, and visual quality in different ways. We study this question through a systematic comparison of prominent test-time strategies including sampler-depth scaling, breadth-oriented parallel sampling, step-by-step construction, iterative refinement, and hybrid breadth--depth for compositional text-to-image generation. Our results show that the preferred strategy depends on prompt complexity and cost assumptions: parallel sampling remains competitive for simpler prompts and high-throughput settings, while iterative refinement becomes increasingly useful as prompts require more object, spatial, numeric, and relational bindings. Surprisingly, step-by-step generation does not consistently inherit the benefits of step-by-step reasoning in language models, since early visual decisions about field of view, layout, scale, and resolution can constrain later edits. Across fixed budgets, hybrid policies often provide the strongest accuracy--cost tradeoff by combining parallel exploration with sequential correction. Finally, we show that refinement-based scaling is bottlenecked by verifier and editor reliability, and that adaptive routing can recover much of the benefit of fixed hybrid policies at lower average cost. Overall, our findings show that effective test-time scaling depends on matching the policy to the prompt and deployment bottleneck: breadth supports exploration and throughput, depth supports semantic repair, and hybrid/adaptive policies offer the strongest tradeoff when compositional correctness is the priority.