Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning
Abstract
Test-time reinforcement learning can adapt vision-language models (VLMs) to unlabeled target data, but its effectiveness depends on reliable self-generated learning signals. We analyze consensus-based VLM test-time reinforcement learning across VQA datasets and model sizes and identify two limitations. First, gains from consensus-based test-time training largely come from answer normalization rather than content correction. Second, many initial VLM responses are incorrect due to the model’s limited ability to jointly use the image and the question; consensus rewards derived from these outputs may preserve the grounding errors rather than correct them. We propose TTIQ, a test-time reinforcement learning framework that harnesses image-question dependence for better VLM adaptation. For each sampled response, TTIQ compares token likelihoods under the original image–question pair and image- or question-ablated variants to estimate dependence on each input. It combines these signals with calibrated confidence to construct a response-level reward for jointly grounded responses, while using token-level dependence for fine-grained positive credit assignment. This shifts supervision from response popularity toward grounded, confident answers. Across eight VQA datasets and multiple VLM sizes, TTIQ achieves the best average performance at every scale, generalizes across VLM families, and improves performance on unseen datasets.