Escape the Context Manifold: Preventing DiT In-Context Editing via Conditional Flow Hijacking
Abstract
Diffusion Transformers enable powerful in-context image editing from a reference image and a text prompt, but also make it easier to maliciously edit public images while preserving the private identity of the reference. Existing DiT-oriented protections, such as DeContext, mainly suppresses context-to-target attention to weaken reference utilization. However, this strategy can be insufficient for complex edits where prompt semantics already dominate generation. Instead, we study proactive protection for DiT-based editors from a conditional-flow perspective, and use a local geometric interpretation of in-context editing, where the denoising trajectory is empirically characterized as being steered by prompt-consistent and reference-consistent directions. Thus, effective protection should not simply corrupt the output, but should drive it away from the original context manifold. To this end, we propose Conditional Flow Hijacking, which adds imperceptible perturbations to the reference image. Instead of suppressing context attention, our method preserves the context pathway, but redirects the context-induced steering signals in the intermediate transformer layers, causing the final trajectory to drift away from the reference consistent solution region. Experiments on FLUX.1-Kontext with different datasets and prompts including attribute and complex scene editing show that our method significantly suppresses reference identity leakage while maintaining image quality and prompt alignment, outperforming prior protection baselines, especially under complex scene editing prompts. We further verify the robustness of the proposed protection method as well as its transferability across other models, highlighting the need for proactive defenses for in-context generative models.