Adaptive Internal Readout for Native Multimodal Models
Abstract
Native multimodal models build rich internal representations across visual, multimodal, and language streams, yet their answers are usually read only from the final layer. We study this mismatch between where task information is available and where the model is allowed to read from. We formalize it with the availability--accessibility gap (AAG), whose primary form compares a small-budget internal readout reference with a learned final-only parity head under the same supervised answer space and loss. Across document, chart, diagram, counting, compositional, and broad multimodal reasoning tasks, we find that task-relevant information often appears at internal sites before it becomes accessible to the final output. We then introduce adaptive internal readout (AIR), a small policy that selects a few internal sites from the full model stack. Under matched budgets, AIR closes much of the gap and outperforms static fusion, dense fusion, visual-only readout, and random site controls. Ablations show that the selected sites are functionally important for the learned readout pathway.