Safety Is Not Free: Numerical Format Is a Deployment-Reliability Variable for Vision-Language Models
Bruce C Xu ⋅ Lan Wu
Abstract
Vision-language models (VLMs) are increasingly deployed under 4-bit quantization to meet memory and inference-cost budgets, and this compression is usually qualified by task accuracy alone. Real-world deployment, however, depends on reliability signals a scalar accuracy score cannot see: whether the model still refuses harmful requests, whether that refusal survives a typographic-image attack, and how confident the audit is. We test whether changing only the numerical format of a VLM silently alters its joint capability and safety behavior. Across five open VLMs, we vary only the inference format among FP16, microscaling FP4 (MXFP4), and NVFP4, then jointly evaluate capability on MMBench-V11 ($n=1292$) and refusal behavior on MM-SafetyBench ($n=624$ per cell, with $n=1248$ targeted reruns). Every compressed cell degrades at least one axis. MXFP4 produces three architecture-dependent failure modes: under-refusal with capability loss on InternVL3, over-refusal with capability collapse on Idefics3-Llama3, and capability-only loss on LLaVA-OneVision and Idefics2. Component ablations localize the Idefics3 refusal shift to the language-model path while its capability collapse is primarily vision-side; an independent Llama-3.2-Vision control rules out a generic Llama3 explanation. NVFP4 is consistently milder. A two-judge capacity check preserves the findings, while a sample-size sweep shows that a 208-record canary overstates the affected model set relative to a 624-record audit. For real-world VLM deployment, numerical format is part of the model artifact: compression must be qualified jointly across capability, safety, modality, and uncertainty rather than through a scalar accuracy score.
Chat is not available.
Successful Page Load