The Boundary Falls Inside the Token: A Granularity Bound on Stopping Rules for Occluded Text
Abstract
When a library label covers the end of a printed title, a vision-language model does not stop reading: it finishes the word from its language prior. The natural remedies - stop early, abstain, truncate - all decide where to cut a generated string, and they share a ceiling no tuning removes. On 2,408 book spines whose visible characters are annotated individually, an oracle cutting at the annotated boundary removes all 183 over-extensions yet leaves only 59 correct: the boundary falls inside a subword token, so a cut before it discards real text and a cut after it keeps invented text. The ceiling belongs to segmentation rather than to one model: across eleven tokenizers and four English text domains the boundary falls inside a token 51-81% of the time, so a seam is usually not available where it is needed, and the rate tracks how finely a tokenizer cuts rather than how large its vocabulary is. Nor can the boundary be recovered from the frozen model: six probes fail, the sharpest showing its probability of closing the string is inverted, preferring the fabricated endpoint. It can box the label and still reads through it. Four interventions supplying the boundary from outside each look effective and each collapses under its own control - a box drawn in the wrong place works as well as one drawn in the right place. Only a plain supervised adapter survives held-out, strictly scored evaluation, with limits we state. Finally we price the behaviour we penalise: 83% of a resolvable sample add correct characters, so the case for stopping rests not on the guesses being wrong but on nothing in the output saying which ones are.