The Locality Cost of Semantic Patch Self-Distillation
Abstract
Masked patch self-distillation has become a default ingredient of modern joint-embedding pretraining, and is widely regarded as a uniform enhancement of dense visual representations. In this work, we revisit this assumption and argue that its effect is more nuanced. Through controlled experiments across multiple joint-embedding objectives and model scales, we observe a consistent tension: while the auxiliary patch objective strengthens semantic abstraction, it systematically erodes fine-grained spatial correspondence. To examine this phenomenon beyond the reach of human-annotated benchmarks, we introduce annotation-free diagnostics that probe how much local visual evidence remains accessible in the learned representation. We further provide a population-level theoretical analysis showing that this trade-off is not incidental but follows directly from the conditional nature of masked semantic prediction, under which target variation unpredictable from visible context is provably averaged out. Taken together, our findings recast masked patch self-distillation as an explicit trade-off between semantic abstraction and local detail, rather than a uniform improvement to dense features.