Agree to Disagree: Multimodal Autonomous Negotiation and Calibration for Entity Representation Learning
Abstract
Learning high-quality multimodal entity representations is essential for advancing reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, rigid alignment in existing fusion strategies can bias representations toward a dominant anchor modality, weakening fine-grained complementary cues and propagating noisy signals into relational reasoning. To address these limitations, we propose MANoR, a Multimodal Autonomous Negotiation framework for calibrated Representation learning. MANoR treats modalities as autonomous semantic agents that negotiate before fusion. Specifically, MANoR preserves modality-specific semantics through centroid-guided subspace contraction, which reduces unstable centroid-relative dispersion while retaining entity-level residual semantics. On this stabilized basis, negotiated attention enables selective cross-modal interaction without collapsing modality boundaries. As modalities can still differ in reliability after interaction, uncertainty calibrated fusion estimates dimension-level aleatoric uncertainty and modality-level reliability to compute relation-aware fusion weights. In this way, MANoR couples modality autonomy, selective interaction, and calibrated integration within a unified representation learning framework. Experiments on MKG-W, MKG-Y, TIVA, and KVC16K show that MANoR consistently outperforms strong MMKGC baselines, achieving a relative Hits@1 gain of up to 17.72% on KVC16K.