SAGE: Semantically Disentangled Representation Learning through Latent Geometry Constraint and Large Language Model
Abstract
Disentangled representation learning (DRL) aims to identify and decompose the interpretable underlying factors of observations. However, most DRL methods rely on independence-oriented regularization to pursue statistical factorization, leaving learned factors without explicit semantic grounding. Such independence-oriented regularization induces an intrinsic trade-off between disentanglement and reconstruction. To address these issues, we establish a spectral learnability threshold for independence-oriented DRL, theoretically revealing a mismatch between the factors favored by the learning objective and human-perceivable semantics. We further introduce SAGE, a semantic-guided DRL framework that leverages MLLMs to automatically extract semantic signals from unlabeled data and uses a triangular latent-geometry constraint to anchor each semantic factor to a designated latent dimension, enabling factor-wise control over human-perceivable semantic variations in generated images. Extensive experiments support our theoretical analysis and demonstrate that SAGE achieves state-of-the-art performance across multiple benchmarks. In addition, our framework exhibits high flexibility and can be integrated with most advanced generative models.