Boundary Mass in Feature Space as a Label-Budget Diagnostic for Representations
Abstract
Low-label transfer is often decided by a practical question: after a representation looks useful, how many labels are still needed to learn the downstream boundary? Existing transferability and probe-quality scores rank global alignment, likelihood, margins, or calibration, but they do not measure how much target data remains close to the fitted decision boundary, where scarce labels are most costly. We propose effective boundary complexity, a label-budget diagnostic for frozen representations and linear probes. It measures mass in a thin strip around the learned boundary and normalizes it by class balance, so dense ambiguous regions and rare-class difficulty are both counted. We prove that, under a smooth representation map, this boundary mass does not change because of generic volume expansion: density change and tangential surface change cancel at first order, leaving only stretch or compression in the boundary-normal direction. This yields a cached-feature estimator that fits a probe on one split, counts held-out features near its boundary, and avoids high-dimensional density estimation. Experiments support the diagnostic at three levels. Controlled transformations match the predicted law across 74 maps with 0.38% median relative error. On real frozen features, normal-direction changes move both measured boundary mass and label demand, while tangent changes do not. A differentiable surrogate reduces measured complexity by 16–36% at matched top-1 utility. Finally, CIFAR-100 and Tiny-ImageNet audits show boundary mass acting as a second-stage diagnostic for near-tied representations, with external multiclass results showing how label-grid saturation limits policy-level gains on label grids where many tasks tie.