ASVQ: What Reparameterization Is Just Enough for Efficient Codebook Learning?
Abstract
Addressing codebook collapse in vector quantization models is crucial for efficient discrete representation learning. Recent solutions increasingly rely on expressive codebook reparameterizations, improving adaptability at the cost of additional optimization and computational overhead. This motivates a simple question: what reparameterization scheme is just enough for efficient codebook learning? In this work, we propose Adaptive-Scale Vector Quantization (ASVQ), a lightweight reparameterization method for learning the codebook in a decoupled way. Experiments across image and audio tokenization tasks demonstrate that ASVQ consistently improves codebook utilization, reconstruction quality, and training stability, while remaining competitive with more complex reparameterization-based methods. These results suggest that ASVQ provides a favorable trade-off between codebook expressiveness and efficiency.