AURA: An Autonomous Retouching Agent with Photographic Visual Thinking
Abstract
Professional image retouching is a sequential and reasoning-driven process. Photographers first analyze scene semantics, lighting structure, and depth cues to design a retouching plan, and then progressively apply global tone grading and mask-guided local corrections based on the evolving visual outcome. Recent MLLM-empowered retouching agents either predict global parameters or respond to user-provided retouching instructions. Although showing impressive results, they lack the capacity for autonomous sequential reasoning and fall short of fine-grained local control. We propose AURA, an AUtonomous Retouching Agent that emulates the workflow of expert photographers. Given an input photograph, AURA autonomously assesses the scene and proceeds progressively. At each step, it reasons about what to retouch, executes the adjustment, and feeds the result back as visual context for next-step reasoning, forming a fine-grained visual chain-of-thought that evolves with the image. To support AURA, we contribute the first expert-annotated long-horizon retouching trajectory dataset, where expert photographers edit the input image from scratch with deliberate artistic intent. Based on this dataset, we train AURA in two stages: supervised fine-tuning to learn the retouching trajectory, followed by GRPO-retouching, an agentic reinforcement learning stage that optimizes for perceptual quality and tool accuracy. Our experiments demonstrate that AURA consistently outperforms existing methods in both quantitative metrics and user studies, enabling one-click professional retouching and democratizing masterpiece previously accessible only to skilled experts.