Correction Space Steering for Hallucination Mitigation in Large Vision-Language Models
Abstract
Large Vision-Language Models (LVLMs) are prone to hallucinations, producing responses that are inconsistent with visual inputs. While inference-time activation steering provides a lightweight solution without retraining, existing methods are either limited by coarse global steering vectors or rely on unconstrained full-dimensional vector prediction, which can produce unreliable vector for unseen hallucination patterns. In this paper, we propose \textbf{Correction Space Steering} (CSS), an robust activation steering framework that predicts correction coordinates in a learned low-dimensional correction space. We first conduct an empirical analysis of hallucination-related activation differences and find that they are predictive of hallucination types, form type-aware geometric structures, and exhibit low effective ranks in most layers. We further provide a theoretical analysis showing that these activation differences can be approximated by a low-dimensional subspace. Based on this insight, CSS learns a low-dimensional correction space from activation differences and trains an online predictor to infer input-specific correction coordinates from inference-time hidden states. The resulting correction vectors are then reconstructed within the correction space and applied to steering activations for hallucination mitigation. Experiments on standard hallucination benchmarks show that CSS effectively reduces hallucinations without retraining LVLMs and outperforms existing inference-time mitigation methods.