OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
Abstract
Omni-modal reasoning is pivotal to realizing artificial general intelligence, yet its advancement is critically constrained by the limited availability of large-scale, human-annotated data for complex reasoning. Motivated by this, we propose OmniJigsaw, a self-supervised reinforcement learning framework built upon a temporal reordering proxy task. Centered on the chronological reconstruction of shuffled audio-visual clips, we successively investigate three modality orchestration strategies: (i) Joint Modality Integration (JMI), which retains the complete visual and auditory streams; (ii) Sample-level Modality Selection (SMS), which selects the dominant modality through a global decision mechanism; and (iii) Clip-level Modality Masking (CMM), which adaptively masks modalities at the clip granularity. Our analysis reveals a ``bi-modal shortcut phenomenon'' in JMI and demonstrates that fine-grained CMM mitigates this issue while outperforming SMS. Incorporating lightweight puzzle-quality curation and verifiable rewards, OmniJigsaw yields substantial gains across 15 video, audio, and omni-modal benchmarks, with CMM achieving the strongest overall performance, validating its effectiveness as a self-supervised paradigm for omni-modal learning.