Confidence-Aware Early Exit for Resource-Constrained Multimodal Question Answering
Kaustubha V
Abstract
Modern vision-language models can incur substantial computational, latency, and memory costs because they typically process complete visual and textual context before predicting. This is inefficient when early evidence is already sufficient, particularly in interactive and resource-constrained settings. We investigate confidence-aware adaptive inference for multimodal question answering, where a model decides both what answer to predict and when enough progressively available evidence has been observed. Unlike early-exit approaches that mainly vary network depth or token computation, our formulation adapts the number of multimodal evidence-processing steps. At evidence step k, a CLIP-style image encoder produces a visual representation that is computed once and cached, while a lightweight text encoder represents the growing evidence prefix. The modalities are combined using input-dependent gated fusion, followed by an answer head. Temperature scaling, fitted on held-out validation data, calibrates the maximum predicted probability. If calibrated confidence exceeds a threshold $\tau$, inference terminates; otherwise, the next evidence item is acquired. Intermediate supervision encourages useful predictions from partial evidence. This avoids repeated image encoding and restricts additional computation to the growing text prefix, fusion module, and answer head. We evaluate progressive-evidence versions of ScienceQA, VizWiz, TextVQA, and AI2D. For controlled evaluation, examples are converted by rule-based segmentation into three simulated stages: an early broad description, an intermediate observation, and a final detailed observation. The calibrated policy processes 1.8 of 3 evidence stages on average and, under the same experimental pipeline, reports a 38% relative latency reduction with 84.4% accuracy, compared with 85.2% for static full-context inference. Calibration improves Expected Calibration Error from 0.11 to 0.07 and reduces premature termination. These findings support confidence-calibrated adaptive inference as a promising foundation for deployment-aware multimodal AI under constrained computational budgets.
Chat is not available.
Successful Page Load