Internalize External Competence for Visual Instruction Editing
Abstract
What if you could edit an image simply by drawing on it---circling an object, writing a short label, or sketching an arrow---with minimal or even no typed prompt? We first reveal a surprising capability: when a text-driven image editor is paired with a strong Vision-Language Model (VLM) at inference time, the combined system can directly interpret visual instructions embedded in the image itself, such as on-image text, bounding boxes, and directional arrows. However, this pipeline incurs additional VLM-planning latency, introduces brittle cross-model error propagation, and depends on external APIs or auxiliary model weights. To address these limitations, we introduce Siphon, a framework that internalizes this externally elicited visual-instruction-following competence into a single diffusion-based editor. Rather than relying on an external VLM at inference time, Siphon first uses a VLM planner to synthesize paired visual-instruction supervision, and then transfers this competence into the editor through lightweight LoRA fine-tuning. The resulting model reads, grounds, and executes on-image annotations directly from pixels, while keeping the diffusion sampling cost close to that of the base editor. Across multiple diffusion architectures, Siphon substantially improves spatial controllability, instruction adherence, and marker removal over text-driven baselines, while matching VLM-assisted pipelines without their additional planning latency, API dependence, and cross-model fragility. We do not position rendered long text as a replacement for conventional text prompts; rather, on-image text is mainly intended for short labels and spatially grounded edit intents, while long or complex instructions can still be provided through the standard text channel.