FusionAudit: Pathway-Conditioned Robustness Auditing for Native Multimodal Models
Abstract
As vision-language models (VLMs) transition from frozen-encoder pipelines to native multimodal architectures, the visual channel increasingly serves as both a perceptual input and a carrier of rendered text that competes with user prompts. Current single-attack, pixel-centric evaluations fail to capture this complexity: they collapse distinct visual attack surfaces into a single scalar and ignore how architectural shifts alter vulnerability. We introduce \textbf{PaRMA}, an evidence-based auditing framework that defines pathway-conditioned risk over two primary image-level interventions: pixel perturbation and rendered-text insertion. Instead of arbitrary attack enumeration, PaRMA uses targeted diagnostics to expose architecture-specific blind spots. Notably, we uncover a \textit{see-but-not-deceived} signature where native VLMs attend heavily to typographic overlays but visually reject them, and we identify severe gradient dispersion in deep-fusion models. To exploit these findings, we instantiate \textbf{FusionAudit}, introducing two targeted attacks: \textbf{MANTIS} (margin-aware pixel perturbation) and \textbf{Nat-Evid} (host-congruent semantic insertion). Evaluating 18 open-weight VLMs across a novel nativeness taxonomy, we demonstrate that legacy single-attack rankings (e.g., PGD) are highly unstable for native models. Our diagnostic-driven attacks substantially raise the estimated residual risk, proving that PaRMA provides a falsifiable, evidence-driven auditing methodology to uncover vulnerabilities systematically hidden by standard benchmarks.