M3-HNTM: Hyperspherical Multimodal Topic Modeling with Symbolic and Contextual Evidence
Abstract
Multimodal evidence can improve topic discovery, but dense auxiliary signals can make topics harder to interpret when they replace readable descriptors. This creates an evidence-role problem: symbolic signals should define what topics mean to a reader, while dense aligned signals should help infer which topics a document expresses. We introduce M3-HNTM, a hyperspherical multimodal neural topic model that separates these roles through a shared document-level latent direction. M3-HNTM utilizes a shared von Mises-Fisher (vMF) posterior for document-level topic inference and vMF-mixture symbolic decoders that represent topics as concentration-aware semantic regions. Text Bag-of-Words and speech Bag-of-Acoustic-Words provide reconstructable symbolic evidence, while aligned image-text embeddings enter posterior inference as contextual evidence and help preserve semantic structure. The model applies to image-text, speech-text, and image-speech-text corpora while keeping topics readable through symbolic descriptors. Experiments on five datasets show that M3-HNTM improves the coherence-diversity quality score over strong text-only, hyperspherical, and multimodal baselines. On SpokenCOCO-Tri, the full tri-modal model outperforms both image-text and speech-text variants, indicating that visual and acoustic signals provide complementary evidence for interpretable topic discovery.