T2V-AttnDisrupt: Inducing Hallucinations in LVLMs via Misrouting Visual Evidence Retrieval
Abstract
Large vision-language models (LVLMs) ground language generation on visual content through the text-to-image slice of self-attention, which acts as a prompt-conditioned router for selecting visual evidence. To induce hallucinations in LVLMs, existing adversarial attacks largely operate at the two ends of this pipeline, either perturbing the front-end visual representation or optimizing against output-token objectives. This leaves the intermediate evidence-selection step as an underexplored yet low-cost attack surface, since it does not require heavy semantic manipulation of image features or extensive output-token optimization. To exploit this attack surface, Text-to-Visual Attention Disruption (T2V-AttnDisrupt) is proposed as an untargeted adversarial attack that induces hallucination by optimizing the input image to distort the text-to-image attention distribution relative to its clean reference, without relying on output-tokens. Experiments on four open-source LVLMs show that T2V-AttnDisrupt increases hallucination rates on captioning benchmarks and reduces VQA accuracy, while preserving overall response quality. It transfers across surrogate–target pairs and generalizes from a single captioning prompt to unseen VQA questions. Moreover, it remains effective against representative defenses, including encoder-level robustness, alignment-based fine-tuning, decoding-time hallucination mitigation, and attention-level interventions, indicating that text-to-image attention is still insufficiently protected. Mechanism analyses show that restoring the clean attention pattern largely recovers visual faithfulness, even when the adversarial image is kept fixed, indicating that hallucination can be driven by corrupted evidence routing rather than feature-level corruption. This low-cost and largely unprotected routing step calls for defenses that explicitly safeguard prompt-conditioned text-to-image attention.