Revisiting The Power of Closed-Form: Robust Deep Image Prototype Discovery via Scale Mixtures
Zhikang Xu ⋅ Jiarui Xing ⋅ Jian Wang
Abstract
We present a robust deep clustering framework for unsupervised visual discovery that bridges the gap between statistical robustness and mathematical tractability. Current generative approaches typically rely on either noise-prone Gaussian priors or computationally intensive diffusion models; however, neither paradigm simultaneously offers explicit latent structures and exact analytic optimization. While diffusion models generate strong priors, their reliance on iterative sampling and intractable likelihoods limits their scalability and interpretability in clustering tasks. Our approach addresses these limitations by constructing a latent space governed by a Student's $t$ distribution via a Bayesian scale-mixture formulation. This yields the first fully closed-form augmented evidence lower bound (ELBO) for this model class, effectively eliminating the biased variational approximations and high computational costs inherent in prior heavy-tailed or diffusion-based frameworks. By treating precision as a latent variable, our objective automatically modulates the influence of each data point; it downweights outliers to learn a data-dependent measure of trust. This mechanism leads to superior robustness and more stable manifold learning compared to state-of-the-art approaches. We demonstrate the resilience of our framework on standard vision benchmarks and complex neuroimaging tasks. In real-world brain magnetic resonance imaging (MRI)-driven neurodegenerative disease analysis, our model successfully recovers clean and anatomically coherent clusters where clinical noise typically collapses rare subtypes. These results underscore the ability of our framework to preserve clinically crucial morphological subtypes, allowing precise discovery within clinical interventions.
Chat is not available.
Successful Page Load