Data-Free Metrics Are Not Invariant Under Functionality-Preserving Reparametrisations
Abstract
Data-free methods for analysing and understanding the layers of neural networks offer many metrics for quantifying notions of strong' versusweak' layers, with the promise of increased interpretability. In particular, random matrix theory (RMT) offers data-free metrics that claim predictive power over the quality of pre-trained models at a layerwise precision and the ability to identify model pathologies, indicating a unique relationship with generalisation. As a result, metrics from RMT are championed for pre-training optimisation strategies and post-training compression. We establish that RMT-based metrics are unrelated to performance or training by showing that functionally indistinguishable reparametrisations of a pre-trained model can have arbitrary metrics. We show this across a range of architectures and scales. To create functionally indistinguishable reparametrisations, we exploit the well-established phenomenon of criticality: some layers can be re-initialised or re-randomised without affecting the functional behaviour of the model -- they are called robust -- while others cannot -- they are called critical. Re-initialising or re-randomising robust layers provides functionally indistinguishable reparametrisations for which RMT-based metrics are arbitrary. Moreover, we show that relationships between metrics offered by RMT are spuriously related to generalisation. We conclude by showing that if many pre-trained models have data-free metrics in a `good' range, it is, in part, dependent on model initialisation.