Towards Realistic Conversational Multimodal Instruction Following
Abstract
Instruction following is a core capability of modern foundation models, underpinning assistant-style interaction, agentic workflows, visual understanding, and multimodal reasoning. Existing instruction-following data and methods, however, remain substantially misaligned with realistic user behavior: users often converse with models over multiple turns, provide interleaved images and videos, switch modalities, or implicitly expect earlier constraints to persist. To bridge this gap, we first identify and formalize six core challenges for realistic conversational multimodal instruction following: Constraint-Guided Visual Reasoning, Cross-Modal Context Transfer, Persistent Instruction Tracking, User-Aware Context Modeling, Iterative Grounded Response Refinement, and Visual Consistency Alignment. Driven by these challenges, we develop a multi-agent data construction pipeline that generates diverse, constraint-rich, image- and video-grounded conversations, followed by human verification. We further improve model behavior through a rubric-guided multi-stage training pipeline that decomposes complex instructions into atomic supervision signals and trains models from basic visual constraint following to advanced multi-turn conversational constraint satisfaction. Finally, we establish Multimodal MultiChallenge (MMMC) as a systematic benchmark for realistic conversational multimodal instruction following. The results show that realistic conversational multimodal instruction following remains challenging for frontier models, while targeted rubric-guided post-training improves constraint satisfaction without sacrificing general multimodal ability.