LOCO: Local Light-Aware Object Compositing with Spatially Varying Illumination-Augmented Data
Abstract
In natural scenes, light sources and occlusions from scene geometry create spatially varying illumination, including brightness gradients, cast shadows, and local color shifts that shape how every object appears. A less explored aspect of diffusion-based object compositing is whether an inserted object can adapt to this local illumination at the desired location. The core challenge to achieve this is the scarcity of large-scale training data capturing how object appearance should adapt to location-dependent illumination. We present LOCO, a data augmentation framework for LOcal light-aware object COmpositing that leverages publicly available monocular videos as illumination sources. We treat each video frame as a virtual spotlight anchored at its camera pose, so that viewpoint variation across a video yields diverse light directions. With controllable beam shape, color, and intensity, this design enables large-scale generation of training data for placement-dependent illumination on the fly. We also introduce LOCO-bench, a benchmark of physically captured composites under spatially varying illumination, where the same object is placed at multiple locations within each scene. Experiments on Laval Indoor SV HDR and LOCO-bench show that LOCO outperforms prior methods in local illumination adaptation while preserving object identity.