Do Glimpse Policies See Like Humans? A Behavioral Audit Reveals Dissociated Viewing Priors in Classification-Trained Active Vision
Abstract
Classification-trained sequential glimpse policies have attained competitive performance and efficiency on large-scale recognition benchmarks, and are motivated by analogy to the saccadic architecture of biological active vision. Yet whether these policies reproduce the behavioral priors that organize human free viewing, or merely approximate their spatial distribution as a byproduct of classification reward, has not been systematically examined. We present a diagnostic behavioral audit of AdaptiveNN across three free-viewing datasets, evaluating five dimensions associated with known properties of biological gaze: spatial priors, face prioritization, local return, saccade amplitude structure, and semantic guidance. We show that AdaptiveNN exhibits: (i) spatial behavior that falls below a simple center-bias baseline; (ii) scanpath geometry unlike human scanpaths in amplitude and sequential structure; (iii) revisitation dynamics inconsistent with human free viewing and not reducible to step-size geometry; and (iv) selectively attenuated semantic guidance, with face-specific prioritization center-explained, social sampling attenuated, and object-level rates preserved. In AdaptiveNN, this dissociation suggests that classification reward alone cannot reproduce several behavioral priors organizing human free viewing, even when object-level sampling is preserved. These findings identify which behavioral priors require constraints beyond classification utility, offering targets for designing cognitively aligned glimpse policies.