Learning to Group and Order: Cross-Instance Self-Supervised RL for Vision-Centric MLLMs
Abstract
Reinforcement learning (RL) post-training has become an important paradigm for improving multimodal large language models (MLLMs), especially with automatically verifiable rewards. For vision-centric MLLMs, self-supervised visual pretext tasks provide such rewards without human annotations. However, existing self-supervised RL methods mostly exploit within-instance structure, such as spatial or temporal ordering within a single image or video, offering limited supervision for cross-instance discrimination and fine-grained comparison. We introduce Group-and-Order Self-Supervised Reinforcement Learning (GO-SSL), a cross-instance jigsaw framework for vision-centric MLLM post-training. Given two visual instances, GO-SSL mixes their local elements and trains the model to group them by instance of origin while recovering the order within each group, coupling instance-level comparison with spatial, temporal, or geometric structure recovery. This paired formulation turns the same unlabeled data into richer verifiable supervision through diverse pairings, shuffles, and hard-pair curricula. We instantiate GO-SSL mainly on cross-image jigsaw tasks and further extend it to 3D depth and video temporal ordering. Across diverse benchmarks, GO-SSL consistently improves over the base MLLM and prior baseline methods, with clear gains in fine-grained perception and spatial understanding. These results suggest a data-efficient direction for self-supervised RL in MLLMs: constructing comparative cross-instance contexts can provide richer transferable supervision than merely scaling data volume or adding isolated pretext tasks. Code will be released later.