Depth-Dependent Safety Geometry in Vision-Language Models: Layer-Wise Localization and Bidirectional Refusal Control
Abstract
Vision-Language Models (VLMs) are widely deployed in multimodal systems, yet the internal mechanisms governing their safety behavior remain poorly understood. We introduce a lightweight inference-time framework for localizing and manipulating safety-sensitive representations without retraining or modifying model parameters. On Qwen3-VL-4B, activation divergence analysis reveals non-uniform safety geometry across depth. Causal interventions show that early vision layers support strong bidirectional refusal control through directional steering, with unsafe rejection spanning 6\%--96\%, whereas later layers are more responsive to activation magnitude scaling and exhibit a substantial safety--selectivity trade-off. We further observe qualitatively similar depth-dependent activation geometry in InternVL2.5-4B despite differences in precise layer location and magnitude. Together, these results distinguish representational separability from causal controllability, suggest that depth-dependent safety geometry recurs across VLM architectures, and expose a white-box representation-level attack surface that motivates inference-time stabilization of safety-sensitive representations.