Computational Depth Predicts Quantization Sensitivity in Multimodal Models
Dongnan Gui ⋅ Ruizhe Wang
Abstract
Post-training quantization degrades multimodal model capabilities unevenly, yet practitioners lack a principled way to predict which capabilities will collapse. We propose the **Computational Depth Principle** (CDP): end-to-end fidelity under quantization follows $\hat p^n$, where $n$ counts sequential precision-dependent operations and $\hat p$ is a per-architecture survival rate. Across four VLM families and diverse benchmarks spanning image understanding, image generation, video understanding, and video generation, we show that deeper reasoning tasks consistently degrade more steeply and that a single shallow probe predicts the full degradation ranking on held-out architectures. The same exponential form extends to image generation, where denoising step count plays the role of depth, and to video understanding, where longer temporal reasoning chains amplify losses. For video generation, spatial quality degrades more than temporal coherence, indicating that the dominant compounding axis is denoising depth rather than cross-frame coupling. A complementary distortion-to-information (D/I) flip ratio diagnoses quantization method effects, revealing that certain methods preserve near-baseline quality while others produce catastrophic collapses on specific architecture pairings that aggregate accuracy alone cannot detect. Module ablation localizes the sensitivity to the language backbone rather than the vision encoder. Together, depth and D/I convert full-suite evaluation into a two-probe screening protocol that surfaces the highest-risk capability gaps first.
Chat is not available.
Successful Page Load