DynaSub: Adaptive Subgrouping for Scalable Representation Learning
Abstract
Real-world observational data often contain existing or emerging heterogeneous subpopulations that deviate from global patterns. Without the ability to detect Out-of-Distribution (OOD) data and model underlying structure, systems risk failing to adapt to emerging patterns, leading to inaccurate or potentially harmful predictions. We introduce DynaSub, a Dynamic Subgrouping Variational Autoencoder that couples latent representation learning with subgroup discovery. DynaSub operates on pre-trained foundation models or regular encoders, learning latent embeddings that define subgroup structure and are iteratively refined to enhance subgroup separability and OOD sensitivity. It incorporates a nonparametric clustering mechanism directly in latent space, enabling the number and structure of subgroups to adapt dynamically during training. DynaSub achieves competitive performance on near- and far-OOD detection across ResNet and ViT-based foundation encoders, reducing false positive rates by up to 10\% under covariate shift while maintaining high AUROC across multiple OOD benchmarks, and excelling in class-OOD settings where entire classes are unseen during training.