The Optimization Prior: Instilling depth for shallow networks, detail for coarse networks
Abstract
Initialization determines the optimization ``fate'' of a model, going far beyond just setting the scale of its weights. It establishes a fundamental initial geometry over the data, permanently dictating which examples are considered close, which directions are easy to change, and which representations gradient descent will naturally refine first. Because this starting point dictates the model's future trajectory, we ask: can this initial geometry be explicitly chosen by referencing a completely different architecture? To achieve this, we introduce representational similarity as an optimization prior. This is an initialization-only procedure that aligns a target network to the representational geometry of a randomly initialized guide network before any downstream training occurs. Crucially, the guide network transfers no learned knowledge: it is frozen, never sees labels. We demonstrate that this cross-architecture transfer of fate works in two key domains: instilling the optimization benefits of depth into shallow targets by aligning them to deeper, randomly initialized guides, and transferring representational granularity by using fine-resolution guides to shape the initialization geometry of networks built with coarser representations. Ultimately, this establishes that a model's architectural destiny can be decoupled from its physical structure and explicitly programmed at initialization.