What Sketches Tell Us about LVLMs: Conventions, Grounding, and Localisation
Abstract
LVLM benchmarks tell us whether a model answers correctly, but not which visual conventions it has retained, nor where those conventions live. We use sketches to expose this hidden structure. A sketch removes texture, lighting, and photographic context while preserving the strokes humans judged sufficient for recognition; if an LVLM can operate on that abstraction, the surviving signal is structural rather than merely photometric. We introduce a training-free sketch probe with three readouts: temporal part-construction order, spatial attention-based grounding, and the mechanistic layer/head locus of grounding. We evaluate across nine LVLMs and five LLaVA-architecture variants on two part-annotated sketch datasets. Three findings emerge. First, LVLM part order tracks human convention: models reproduce canonical orderings where human drawings converge, and become correspondingly variable where human orderings vary. Second, selected frozen LVLM attention heads ground object and scene sketches competitively with specialised grounders, exceeding CLIPSeg and GroupViT on mAcc@0.1 despite using no grounding supervision. Third, the grounding signal concentrates in a narrow mid-network band whose location follows the base language model rather than parameter count, and remains stable across object and scene sketches. The result is a practical diagnostic for sketch-driven LVLMs: no fine-tuning, no auxiliary segmenter, and a small architecture-dependent layer band that doubles as a head-selection prior, improving grounding mAcc@0.1 in every model and dataset we tested.