Timezone: »
The majority of existing multimodal sequential learning methods focus on how to obtain effective representations and ignore the importance of multimodal fusion. Bilinear attention network (BAN) is a commonly used fusion method, which leverages tensor operations to associate the features of different modalities. However, BAN has a poor compatibility for more modalities, since the computational complexity of the attention map increases exponentially with the number of modalities. Based on this concern, we propose a new method called generalizable multi-linear attention network (MAN), which can associate as many modalities as possible in linear complexity with hierarchical approximation decomposition (HAD). Besides, considering the fact that softmax attention kernels cannot be decomposed as linear operation directly, we adopt the addition random features (ARF) mechanism to approximate the non-linear softmax functions with enough theoretical analysis. We conduct extensive experiments on four datasets of three tasks (multimodal sentiment analysis, multimodal speaker traits recognition, and video retrieval), the experimental results show that MAN could achieve competitive results compared with the state-of-the-art methods, showcasing the effectiveness of the approximation decomposition and addition random features mechanism.
Author Information
Tao Jin (Zhejiang University)
Zhou Zhao (Zhejiang University)
More from the Same Authors
-
2022 Poster: GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech »
Rongjie Huang · Yi Ren · Jinglin Liu · Chenye Cui · Zhou Zhao -
2022 Poster: Towards Effective Multi-Modal Interchanges in Zero-Resource Sounding Object Localization »
Yang Zhao · Chen Zhang · Haifeng Huang · Haoyuan Li · Zhou Zhao -
2022 Poster: Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-Speech »
Ziyue Jiang · Zhe Su · Zhou Zhao · Qian Yang · Yi Ren · Jinglin Liu · 振辉 叶 -
2022 Poster: M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus »
Lichao Zhang · Ruiqi Li · Shoutong Wang · Liqun Deng · Jinglin Liu · Yi Ren · Jinzheng He · Rongjie Huang · Jieming Zhu · Xiao Chen · Zhou Zhao -
2022 Spotlight: Lightning Talks 4B-4 »
Ziyue Jiang · Zeeshan Khan · Yuxiang Yang · Chenze Shao · Yichong Leng · Zehao Yu · Wenguan Wang · Xian Liu · Zehua Chen · Yang Feng · Qianyi Wu · James Liang · C.V. Jawahar · Junjie Yang · Zhe Su · Songyou Peng · Yufei Xu · Junliang Guo · Michael Niemeyer · Hang Zhou · Zhou Zhao · Makarand Tapaswi · Dongfang Liu · Qian Yang · Torsten Sattler · Yuanqi Du · Haohe Liu · Jing Zhang · Andreas Geiger · Yi Ren · Long Lan · Jiawei Chen · Wayne Wu · Dahua Lin · Dacheng Tao · Xu Tan · Jinglin Liu · Ziwei Liu · 振辉 叶 · Danilo Mandic · Lei He · Xiangyang Li · Tao Qin · sheng zhao · Tie-Yan Liu -
2022 Spotlight: Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for Text-to-Speech »
Ziyue Jiang · Zhe Su · Zhou Zhao · Qian Yang · Yi Ren · Jinglin Liu · 振辉 叶 -
2022 Spotlight: GenerSpeech: Towards Style Transfer for Generalizable Out-Of-Domain Text-to-Speech »
Rongjie Huang · Yi Ren · Jinglin Liu · Chenye Cui · Zhou Zhao -
2022 Spotlight: M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus »
Lichao Zhang · Ruiqi Li · Shoutong Wang · Liqun Deng · Jinglin Liu · Yi Ren · Jinzheng He · Rongjie Huang · Jieming Zhu · Xiao Chen · Zhou Zhao -
2022 Poster: Unsupervised Representation Learning from Pre-trained Diffusion Probabilistic Models »
Zijian Zhang · Zhou Zhao · Zhijie Lin -
2021 Poster: PortaSpeech: Portable and High-Quality Generative Text-to-Speech »
Yi Ren · Jinglin Liu · Zhou Zhao -
2020 Poster: Counterfactual Contrastive Learning for Weakly-Supervised Vision-Language Grounding »
Zhu Zhang · Zhou Zhao · Zhijie Lin · jieming zhu · Xiuqiang He -
2019 Poster: FastSpeech: Fast, Robust and Controllable Text to Speech »
Yi Ren · Yangjun Ruan · Xu Tan · Tao Qin · Sheng Zhao · Zhou Zhao · Tie-Yan Liu -
2018 Poster: MacNet: Transferring Knowledge from Machine Comprehension to Sequence-to-Sequence Models »
Boyuan Pan · Yazheng Yang · Hao Li · Zhou Zhao · Yueting Zhuang · Deng Cai · Xiaofei He