Show Me What You Don’t Know: Efficient Sampling from Invariant Sets for Model Validation
Abstract
The performance of machine learning models depends on discovering features in the data that are relevant to solving the given task while ignoring irrelevant variation. To check and visualize whether this is the case, we propose ‘invariance auditing’, a diagnostic tool to analyze feature extractors. Invariance auditing generates diverse inputs which all yield the same features as a given query input, and are therefore treated as identical by the model. Using external knowledge about the task at hand, this enables assessing if the learned invariances do or do not conform to the task's required semantics. Unlike existing work where a dedicated generative model is trained for each feature extractor, our algorithm is training-free and exploits a pretrained diffusion or flow-matching model to sample invariant inputs. Our fiber loss -- which penalizes feature mismatch -- guides the denoising process so the output shares the same representation as the query input. This replaces days of training with a single guided generation procedure at the same quality. Experiments on popular datasets and model types demonstrate that our auditing method reveals invariances spanning very desirable and concerning behavior. For instance, it detects cases where Qwen-2B places patients with situs inversus (heart on the right side) in the same fiber as typical anatomy.