Matched-Control Tests of Partition-Source Claims in One Routed Distillation Family
Abstract
Partitioned routed distillation is hard to interpret when route count, supervision grouping, and search or filtering budget change together. We study that attribution problem inside one fixed routed family using two matched nulls: a routing-only control (Latent-Routing MoA) and a structure-matched shuffled-partition control (Matched-Partition-Shuffle), with teacher tokens, trainable parameters, and router depth matched exactly. The decisive result is Table tab:main: the zero-manual deployable row ModeDistill-Auto-v4-Open, whose partition source is induced by GPT-OSS-120B-Instruct pairwise same-strategy judgments (Apache- 2.0 , distinct from the trace-generating teacher), beats MoA by +2.46 pp and Matched-Partition-Shuffle by +2.04 pp on the five-domain primary-shift average (A1), with positive OOD gaps on all five primary domains. Table tab:externalbreadthtransfer is intentionally weaker evidence: for that exact same headline row it reports only four external benchmark aggregate means, where Auto-v4-Open beats MoA by +4.91 / +3.99 / +3.70 / +2.87 pp, and it does not support slice-wise breadth or cross-family generality. The open pipeline remains within 0.05 pp of the manual upper-bound reference on A1 and within the pre-specified 0.50 pp tie band on every external mean, so we keep it as the deployable row and retain the manual row only as calibration. The claim is therefore narrow: within this routed family, once routed budget is matched, the residual OOD gain is attributable to the externally induced partition source rather than to routing alone or to structure-matched shuffled partitions. Figure fig:identification and the appendix are corroborative only. The weakest primary shift remains ALFWorld tool reconfiguration ( +1.15 pp vs.\ MoA, paired CI [-0.12,+2.41] ), and discovery-time verifier calls and teacher generation are disclosed separately from the matched comparison.