Beyond Single-Shot Conditioning: Test-Time Condition Refinement for Diffusion-Based Image Restoration
Abstract
Conditional diffusion models have emerged as a powerful paradigm for image restoration, leveraging pre-trained generative priors to recover high-fidelity details from degraded inputs. However, existing methods typically follow a static inference protocol, treating the degraded observation as a fixed control signal throughout sampling. This overlooks the potential of inference-time scaling, where guidance can be refined on the fly to improve output quality. Recent test-time optimization approaches that manipulate noise variables or denoising trajectories often compromise structural fidelity, while those targeting text embeddings lack the spatial precision needed for restoration. To address this, we propose Condition-Aware Test-Time Optimization (CATTO), a training-free method that iteratively refines the visual conditioning signal at inference time. CATTO performs reward-aligned condition refinement under the pre-trained prior to enhance perceptual quality. To avoid complex backpropagation, we design a gradient-free optimization strategy guided by a joint objective: maximizing a perceptual reward while enforcing trajectory- and condition-level consistency to explicitly preserve structural fidelity. This process is accelerated by optimizing within a low-dimensional frequency subspace and reusing reliable updates across nearby denoising steps. Experiments on image restoration benchmarks show that CATTO improves perceptual quality over strong diffusion-based baselines without updating model parameters.