Local-Interaction Learning Dynamics: A Markov Random Field Framework for Convergence of Deep Neural Network Learning
Wen Dong
Abstract
Modern neural-network training is governed by local interactions: residual streams carry signal, normalization changes scale, attention and state-space blocks route information, augmentation changes which directions generalize, and optimizer state determines the actual checkpoint update. Parameter-space curvature and full-weight posterior views reveal global geometry, but they do not show which layer, block, or interface created it. We introduce a computation-graph view in which parameters $\theta$ and intermediate states $u$ define a local Markov--Gibbs system. A potential energy encodes module relations, data loss, and regularization; its Gibbs law gives a probability reference, and its checkpoint-local Fisher/Gauss--Newton precision matrix gives the curvature geometry. This precision matrix is sparse because variable groups interact only when they share a local module or loss factor. Eliminating intermediate states recovers the reduced parameter curvature. From this common object we prove two results: separator transfers give two-sided criticality certificates for upper stability and active plasticity, and a resolved KL-to-Gibbs bound yields bounded-loss PAC-Bayes certificates on directions explored by training. Checkpoint measurements on CIFAR-10 and GPT-20M/FineWeb-Edu show that these quantities diagnose learning-rate stability, plasticity, and generalization geometry without dense Hessian, Fisher, or covariance matrices.
Chat is not available.
Successful Page Load