Reading the Unreadable: Text-Aware Image Super-Resolution Needs Reasoning
Abstract
Text-aware image super-resolution (TAISR) aims to recover high-resolution images while preserving textual content. Yet some degraded text cannot be restored from local visual evidence alone: humans often read it by reasoning over image context, logical patterns, and prior knowledge. This raises a fundamental question: does TAISR need reasoning? We answer with two new diagnostic tools: ReasonText, a human-curated benchmark with perception-grounded difficulty levels that separate reasoning-required text from locally recoverable text, and GTTCA, a crop-based recognition accuracy metric that isolates restoration quality from detection errors. Evaluating twelve recent SR models, we find a consistent and substantial performance drop on reasoning-required text, revealing a systematic blind spot in current TAISR designs. To address this limitation, we build on two observations: (1) SR models benefit from ground-truth text labels provided as captions, and (2) modern MLLMs can infer degraded text from LR images through reasoning. These observations motivate Reasoning Transfer via Captioning (RTC), a training-free plug-in that uses an external MLLM as a reasoning module and injects its inferred text into SR models through natural-language captions. Surprisingly, with RTC, real-world SR models can even outperform dedicated text-aware SR models, suggesting that reasoning should be treated as a core design axis for future TAISR.