Seeing Through a Slit: Anorthoscopic Perception in Multimodal Large Language Models
Abstract
Human vision can recover the complete structure of a highly occluded object from partial views revealed successively through motion. This phenomenon, known as anorthoscopic perception, relies on the visual system’s ability to integrate information over time to form a perception of the complete object. A classical way to study this phenomenon is slit viewing, where an object moves behind a narrow stationary vertical slit so that only a small portion of it is visible at any instant. In this study, we use slit viewing to examine whether multimodal large language models (MLLMs) can similarly integrate visual information that is distributed across time. We test 50 familiar three-letter English words selected from SUBTLEX-US and 50 matched anagram pseudowords to examine the effect of lexical priors. We generate videos in which these strings move behind a stationary vertical slit spanning 20%, 10%, or 5% of the stimulus width. As the slit becomes narrower, progressively less of the string is visible in any individual frame, although the complete string is revealed over the course of the sequence. We also test the models on fully visible stimuli to establish a baseline, under which nearly all models identify the stimuli at or near ceiling. However, under slit-viewing conditions performance declines substantially as the slit becomes narrower. Even the strongest model shows a marked deficit: at the 5% slit width Gemini-3.5-Flash reaches only 76% accuracy, and several open-source models fall close to 0%. In contrast, human observers identify all stimuli correctly across all slit widths, showing no degradation from the full-view condition to the narrowest slit width. These findings reveal a striking dissociation between recognition and anorthoscopic perception in current MLLMs: stimuli that are readily recognized when fully visible become increasingly difficult to recover as less of the same spatial information is visible in any single frame.