Format Compliance Masquerading as Reasoning Gain in Multimodal RLVR: A Case Study
Guneesh Gupta ⋅ Anshul Tripathi
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become a standard way to improve step-by-step reasoning in language models on problems such as those in mathematics. A growing line of work asks whether the resulting gains reflect genuinely new capability or a reweighting of behavior the model already had. Vision-language models (VLMs) raise this question as well. We apply pass@1 and pass@$k$ analysis to a setting where a VLM base model is trained on text only and analyse how changes carry over when the identical problem is presented as an image. We evaluate under three conditions: (T), where the model solves plain text; (D), where the model transcribes an image to text before solving its own transcription; and (E), where the model reads and solves the image in a single pass. Because our training reward and one natural evaluation convention are built on the same answer extractor, a known risk in RLVR evaluation, we designed our study to test that risk directly, scoring every completion under two conventions throughout (one format-strict, the other format-agnostic). To test the difference between these conventions, we conduct control experiments examining the extent to which a single line added to the base model's prompt can reproduce the gain from RL training. Testing on a text-only dataset outside the training distribution provides evidence that sharing an answer-extraction function between the training reward and the evaluation metric can produce errors in both directions: false positives and false negatives. We further show that format-strict accuracy factors exactly into format compliance and accuracy conditional on compliance, and that under this decomposition the two models are equally accurate among completions that comply, in every condition.
Chat is not available.
Successful Page Load