CASQRec: Collaborative-Adaptive Semantic Quantization for Multimodal Recommendation
Abstract
Multimodal recommendation leverages item content, such as text, images, and audio signals, to complement sparse user-item interactions. Existing methods mainly combine content and collaborative signals at the representation level through graph propagation, feature fusion, or auxiliary alignment. However, continuous item embeddings make semantic sharing implicit and provide limited support for transferring behavioral evidence through reusable semantic factors. Discrete semantic codes offer explicit shared variables, but existing code-based recommenders usually treat item-code assignments as fixed outputs of pretrained quantizers, clustering algorithms, or frozen encoders, making the assignment mechanism insensitive to user behavior. We propose \textbf{CASQRec}, a \textbf{C}ollaborative-\textbf{A}daptive \textbf{S}emantic \textbf{Q}uantization framework for multimodal recommendation. CASQRec learns item-code assignments as content-inferable routing decisions supervised by collaborative behavior during training. It constructs a semantic-code graph, where fixed content-derived RQ codes are graph nodes and item-code edges are predicted from item content by a learnable prior router. To make routing behavior-aware while keeping item-code routing independent of item-side interaction history at inference time, CASQRec introduces a collaborative posterior teacher only during training and distills its assignment distributions into the content prior. The resulting prior-induced user-item-code graph allows collaborative evidence to be shared through discrete semantic codes while keeping item routing content-based. Experiments on three public datasets show that CASQRec consistently outperforms representative multimodal recommendation baselines. The code resources are available in the supplementary materials.