What Happens When the Font Changes? A Controlled Study of Typographic Robustness in Vision- Language Models
Abstract
Vision-Language Models (VLMs) perform well on multimodal reasoning benchmarks, yet their robustness to typographic variation remains largely unexplored. We introduce FontReason, a controlled framework for evaluating reasoning robustness under variations in font family, font size, and realistic visual degradation while holding semantic content fixed. Using chain-of-thought prompting as the primary protocol, we evaluate 3 closed-source and 12 open-source VLMs and 6 extraction–reasoning pipelines on 29,427 samples spanning rendered-text and visually grounded reasoning. In addition, we include direct-answer and program-of-thought prompting as supplementary analyses. Our experiments reveal that stylized fonts generally reduce reasoning accuracy, particularly for smaller models, while visual noise further amplifies these failures. We also observe that extraction–reasoning decomposition substantially improves rendered-text reasoning but underperforms end-to-end models on charts, where extraction can discard important visual structure. These findings establish typography as an important axis of multimodal robustness evaluation.