FreDRec: Frequency-Decoupled Knowledge Distillation for Multimodal Recommendation in Missing Modalities Scenarios
Abstract
Multimodal Recommender Systems (MRSs) have achieved significant success in personalized recommendations by integrating diverse modalities such as images, text, and audio. While in complex real-world industrial scenarios where MRS has usually been deployed, the presence of missing modalities is very common. Most existing MRS methods recover missing data by disentangling available modality features into general and specific representations, thereby generating missing modalities based on personalized aligned feature. However, these methods rely on implicit constraints within pre-extracted multimodal feature spaces, which usually introduce noise that prevents the effective decoupling of general and specific semantics, thereby diminishing recommendation accuracy. Moreover, unguided linear transformations of available modality features for generation lead to semantic hallucinations, which inevitably degrade the ranking quality of the recommendation results. To address these issues, we propose the Frequency-Decoupled Knowledge Distillation Framework for Multimodal Recommendation (\textbf{FreDRec}), which conducts effective and robust missing feature generation. Specifically, motivated by the physical characteristics of multimodal signals, FreDRec utilizes Fourier and Discrete Wavelet Transforms to explicitly decouple raw multimodal signals into low-frequency core semantics and high-frequency details, which is mathematically proven to effectively minimize noise introduction at the source. Next, to ensure robust feature generation, we pre-train a teacher model based on decoupled multi-view graph propagation to guide a student model via multi-level knowledge distillation, which establishes an end-to-end highly efficient lightweight architecture and achieves precise missing modality generation with expert collaborative bounds. Finally, extensive experiments on multiple public benchmark datasets demonstrate that FreDRec achieves superior performance under various modality missing ratios.