Pretext Reasoning: Scaling the Building Blocks of Interleaved Multimodal Reasoning
Abstract
Human reasoning is not purely linguistic: for visual problems, people often think by changing the visual representation itself. Interleaved multimodal reasoning seeks to bring this ability to unified models by allowing them to construct intermediate visual states and use them within multimodal reasoning traces. Yet scaling this ability is difficult, since task-driven instruction tuning must teach visual-state construction, faithfulness, and downstream reasoning all at once. We introduce PRIMER (\textbf{PR}etext-based \textbf{I}nterleaved \textbf{M}ultimodal r\textbf{E}asoning \textbf{R}ecipe), a two-stage training recipe that uses pretext reasoning as a primer for interleaved multimodal reasoning. Stage 1 builds \pretext, an instruction-free corpus that converts classical self-supervised vision pretext tasks into Thought--Image--Thought traces, teaching the model the reusable mechanics of producing and reading visual states. Stage~2 builds \instruct, a task-driven instruction-tuning corpus that teaches task-relevant visual-state construction across perception, spatial understanding, and mental world modeling. On a 7B unified multimodal model, Primer improves over the matched-budget BAGEL reference by +10.40, +5.93, and +15.70 across the three cognitive levels, with shorter traces, and rivals models several times its size on spatial and world-modeling averages. These results establish classical self-supervised pretext tasks, repurposed as Thought-Image-Thought traces, as a scalable, parameter-efficient primer for interleaved multimodal reasoning, the recipe at the core of Primer.