Alignment Without Binding: A Credit-Depth Threshold for Cross-Modal Transfer in Local Learning
Anthony Yeh ⋅ Simon Sang
Abstract
A child learns which words go with which things from a stream of paired experience that is small, ambiguous, and never labelled at the level of the pairing. Work on sample efficiency tests that setting by shrinking the data budget while leaving the learning rule fixed at backpropagation. We vary the rule instead. Local, layer-wise learning rules can acquire transferable features without backpropagation. It is unclear whether this extends to cross-modal binding, where the transferable object is a correspondence between two encoders rather than a feature within one. We train image and caption encoders from scratch on 20,000 pairs and about 208k words, roughly 0.2% of this programme's 100M-word developmental reference, under a symmetric InfoNCE coupling and vary a single knob: the per-layer credit depth $d$, how many blocks a local contrastive gradient reaches before a stop-gradient, from strictly per-layer credit ($d=1$) to backpropagation through the trunk ($d=4$). At 156M, transfer is a step function of $d$. Strictly local credit fits the training coupling at least as well as every other arm (0.99 train latent retrieval, above the backpropagation comparator's 0.955) yet transfers 1.10× held-out category precision@10 over a base rate of 8.1%. One additional layer of reach recovers most of backpropagation's lift and $d=3,4$ plateau. Under a shared evaluation pool and a matched contrastive batch, the gap against a comparator matched in architecture, data and split but not in optimizer, parameterization or objective structure is +0.74, +0.45 and +0.48 at 156M, 330M and 0.70B parameters, resolvable at every width. That gap is not a credit-reach effect: our own full-reach limit tracks $d=2$ to within 0.036 at all three widths. We refute four natural accounts by measurement, three of that deficit and one of the $d=1$ failure, namely a batch cost growing with width, content-selective binding failure, updates converging toward backpropagation's, and benign neglect of the trunk. Two of these measurements are surprising on their own. Gradient agreement with the full-reach limit falls with width while transfer relative to backpropagation improves, and freezing the encoder at its random initialization transfers better than training it under strictly local credit, on three seeds whose range is disjoint from every measurement of the trained arm. We also report three ways to manufacture a locality result, one of which, the evaluation split, accounts for 52% of apparent seed noise. The one intervention that moves the coupling is caption density, which raises lift from 2.26 to 4.35× across three seeds with disjoint ranges, that is, five descriptions of one scene rather than one.
Chat is not available.
Successful Page Load