Training-Free 3D Editing via Feature-Divergence Localization and Trajectory Correction
Abstract
Modifying 3D assets requires enacting targeted structural transformations while strictly preserving the geometric identity of unedited regions. While recent probability flow models enable training-free editing, bounding these modifications within a high-dimensional latent space presents severe mathematical challenges. Current inversion-free trajectories rely on global velocity updates, which consistently result in spatial bleed into preserved structures and induce trajectory under-commitment, causing edits to stall mid-path. To resolve these dual failures, we introduce FocusFlow, a segmentation-free framework that structures the integration trajectory into two disjoint phases. During early coarse structure formation, FocusFlow applies Feature-Divergence Localization (FDL) to spatially gate the velocity field using intrinsic cross-attention. During late-stage fine-detail refinement, Reference-Guided Generation (RGG) drives topological convergence by pulling the trajectory toward an explicit clean-latent target. Evaluated on the joint Google Scanned Objects and PartObjaverse-Tiny benchmarks, FocusFlow breaks the traditional preservation-modification trade-off. It achieves superior structural preservation while maximizing edit magnitude. By providing stable, localized geometric control without manual spatial annotations, this phase-scheduled approach removes significant technical barriers to the precise modification of synthetic 3D media.