Learning When to Think: Dual-Reference Offline Optimization for Adaptive VLM Reasoning
Abstract
Reasoning traces improve vision-language models on complex multimodal tasks, but forcing every input to follow a long reasoning process leads to overthinking and unnecessary inference cost. We study adaptive fast/slow thinking for VLMs, where a single model should answer directly for simple inputs and reason step by step only when needed. The key observation is that offline adaptive thinking is not merely a length-control problem: fast direct answers and slow reasoning traces are typically generated by distinct behavior policies. This dual-reference structure makes standard offline reward optimization, such as Decoupled Generation and Optimization (DGO), mismatched because it assumes a single reference policy and can misweight one mode. We propose \textsc{Dual-Reference DGO}, which models fast and slow responses under a fast/slow mixture reference and derives a reward-weighted offline objective for adaptive VLM reasoning. Our formulation shows that router-free fast/slow switching emerges from competition between fast and slow partition functions, and yields an efficient mode-wise correction for behavior-reference mismatch. Experiments on multiple datasets show that \textsc{Dual-Reference DGO} improves the accuracy--length trade-off over baselines. Ablations show that single-reference DGO tends to collapse to almost always-fast or always-slow behavior, while our dual-reference formulation preserves both modes and learns difficulty-dependent routing.