Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
Abstract
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods tend to outperform contrastive ones, a pattern that sits awkwardly with recent evidence that pixel reconstruction yields uninformative features for natural-image classification. We argue this discrepancy is not an accident, but a consequence of how HTR signal is distributed in pixel space. By probing the input distribution directly, we show that HTR's discriminative content is concentrated in the high-variance pixel subspace and is essentially absent from the low-variance one, the inverse of the structure observed for image classification. Under this view, the right SSL family becomes predictable: objectives that preserve high-variance pixel content should transfer best. We test this prediction across six SSL methods spanning three families (pixel-grounded MIM, JEPA-style, contrastive image--image and image--text), under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. Pixel-grounded SSL produces the lowest CER on every benchmark and every probe, exposes per-position character information that other families recover only via the readout, and is the only family that benefits from real-data pretraining. A geometric property of the encoder, its alignment with the high-variance pixel subspace, predicts CER within every SSL method we test. With a pretrained LLM decoder, a frozen pixel-grounded encoder is already competitive with fully fine-tuned supervised baselines, and full fine-tuning beats them on the mean and ranks first or second on every benchmark. These results challenge the prevailing view that pixel reconstruction wastes capacity on irrelevant detail: whether reconstruction is wasted depends on where the discriminative signal lives in the input.