RealNLR: Do VLMs Learn Visual Algorithms from Synthetic Fine-Tuning?
Abstract
Vision-language models (VLMs) score well on standard benchmarks yet fail at simple tasks that require chaining visual evidence across an image, termed nonlocal visual reasoning (NLR): re-identifying a transformed object (object re-identification), following a chain of cues (saccadic search), and tracing a wire through clutter (contour tracing). We ask whether fine-tuning can teach these visual algorithms and, if so, what the model actually learns. We fine-tune five open VLMs (4B to 32B) with QLoRA on synthetic versions of the three task families, varying which side of the model is adapted (language or vision), whether supervision shows only the answer or every intermediate step, and the amount of data. To measure transfer, we introduce RealNLR¹, a benchmark that poses the same tasks on realistic imagery with matched distractors. We find that (i) fine-tuning raises in-distribution accuracy by 13 points on average, and the best runs go from near zero to 98%, but transfer to RealNLR is zero; (ii) process supervision teaches the procedure on synthetic data, with InternVL3-14B reproducing 96% of out-of-distribution search trajectories exactly, yet the model still reports the answer at the chain length it was trained on; and (iii) internal analyses show that fine-tuning improves only one stage of the NLR procedure, and that adapting the language side helps more than the vision side. Despite near-perfect synthetic accuracy, harder and realistic variants show the models did not learn the algorithm.