Flow-of-Thought: A Framework for Visual Reasoning
Mariia Baidachna ⋅ Nicolas Pugeault
Abstract
Mental imagery, ``seeing with the mind's eye'', enables humans to reason through complex spatial transformations. Yet Large Language Models (LLMs) and Vision Transformers (ViTs) often struggle with tasks that require tracking intermediate visual states. We introduce Flow-of-Thought (FoT), a framework that learns continuous visual reasoning traces using flow matching. By integrating a learned ordinary differential equation (ODE), FoT generates intermediate sketches at any time $t \in [0,1]$. Frozen trajectory fields produce rotation and maze traces, while same-versus-different decisions compare explicit generative hypotheses. We evaluate FoT against a direct ViT, DINOv3 with a trained decision head, Qwen3.5-9B, GPT-4o-mini, GPT-5.6 Sol, Claude Opus 4.6, GPT-4V, and P$^2$ with Qwen3VL-4B. FoT performs strongly on synthetic rotation and maze navigation and shows encouraging transfer to BLINK, supporting continuous visual traces as an effective and interpretable representation for spatial reasoning.
Chat is not available.
Successful Page Load