A Unified Merge Calculus for Learning-Rate Scaling on Neural Computation Graphs
Haosong Zhang ⋅ Wu Shenxi ⋅ Zhiyuan Che ⋅ Xi Chen ⋅ Wei LIN ⋅ Yichi Zhang
Abstract
Learning-rate selection is one of the most consequential and repeatedly tuned decisions in modern deep learning, and the ability to transfer it quickly across widths, depths, and architectures can substantially reduce the cost of scaling new models. Existing transfer laws explain width through maximal-update parameterization and depth through depth-based scaling laws, but many modern architectures are better described as neural computation graphs with heterogeneous merge rules rather than by depth alone. We develop a unified merge calculus for learning-rate scaling on such graphs. Its central object is a graph-aware complexity coefficient that combines local merge semantics with global path structure and induces a $-1/2$ power law for the target learning-rate scale. We further show that depth alone is generally insufficient once merge semantics vary across architecture families. We validate the framework on controlled graph families and lightweight representative architectures covering uniform sum, residual-aware sum, concat, weighted fusion, and hybrid designs. Across these settings, the resulting graph coefficients organize optimal learning rates more effectively than depth alone and support practical learning-rate transfer.
Chat is not available.
Successful Page Load