PAVE: Prefill-Conditioned Activation Editing for Hallucination Mitigation in LVLMs
Abstract
Large vision-language models (LVLMs) can caption images, answer visual questions, and reason over scientific diagrams, yet they still produce hallucinations unsupported by visual evidence. Existing approaches mitigate hallucinations either by reshaping the decoding distribution or by steering hidden activations with calibration-based directions. The former adjusts token probabilities but operates only at the output level, whereas hallucinations often arise when language priors dominate visual evidence in intermediate representations. The latter applies a fixed calibration intervention, which may fail to capture image-specific hallucination drift or inadvertently suppress grounded evidence. To address these limitations, we propose PAVE, a training-free method for input-specific hidden-state intervention. PAVE constructs offline subspaces for hallucination-drift suppression and visual-evidence amplification, and adapts the hidden-state update to each input using a per-image visual basis computed during prefill. By editing hidden states with this prefill-conditioned subspace, PAVE strengthens grounded evidence while suppressing hallucination drift, without relying solely on decoding-level corrections or fixed activation edits. Experiments on LLaVA-1.5-7B and Qwen-VL show that PAVE substantially reduces hallucinations on CHAIR and POPE while preserving grounded utility on MME. Code will be released.