Intrinsic Information Theoretic Analysis of ReLU Nets
Johan Mylius-Kroken ⋅ Elisabeth Wetzer ⋅ Ali Ramezani-Kebrya ⋅ Robert Jenssen ⋅ Kristoffer Wickstrøm
Abstract
Mutual information (MI) is fundamental to representation-learning analysis in deep networks, but in high dimensions MI estimators depend on hyperparameters --- bin widths, kernel bandwidths, neighbourhood radii, auxiliary networks --- whose bias often dominates the signal. We show that continuous piecewise-linear (CPWL) networks, including ReLU networks, admit an intrinsic information-theoretic analysis that avoids this pitfall: their inherent partition of the input space into convex linear regions assigns to every input a discrete region label $\Pi$ alongside the continuous representation $T$. We decompose $I(Y; T)$ exactly into geometric terms, where the \emph{routing information} $I(Y; \Pi)$ admits a hyperparameter-free plug-in estimator that approximates a population lower bound on every other MI term in the decomposition. When the partition saturates, a functional-equivalence quotient at relative Frobenius tolerance $\varepsilon \in [0, 2]$ collapses regions implementing the same linear operator and restores informativeness. Empirically, the routing estimator correlates strongly with three established MI baselines {\it without any hyperparameter}, and the functional quotient at moderate $\varepsilon$ recovers informative estimates in the regime where the raw partition is bounded by $\log_2 N$ rather than $H(Y)$. We provide a parameter-free framework for analysing the internal information geometry of CPWL networks, grounded in their inherent piecewise-linear structure rather than in external discretisation.
Chat is not available.
Successful Page Load