RLVR from Within: Learning to Ask to Elicit Intrinsic Rewards for Test-Time Reasoning
Abstract
Test-time reasoning for large language models (LLMs) often relies on external process reward models (PRMs) or value estimators to guide search over intermediate reasoning states. However, learning such rewards requires assigning credit to intermediate states from final outcomes. We ask whether useful test-time rewards can instead be elicited from the reasoner itself. We propose AskR, which learns what to ask so that the reasoner's own self-verification becomes a reliable intrinsic reward. AskR structures reasoning as sequences of sub-questions and sub-answers, providing dense QA-level self-verification signals without training a separate verifier. Crucially, self-verification alone is not sufficient: a locally verified SubQA need not be predictive of the correct final answer. Rather than learning a new verifier, AskR learns the sub-question policy so that the induced intrinsic reward becomes more predictive of final-answer correctness and therefore useful for test-time optimization. Across our benchmark suite, AskR consistently improves over PRM-guided reasoning under the same optimizer, by 7.12 points on average and up to 9.18 points depending on the optimizer. In its strongest setting, AskR further outperforms RLVR by 2.58 points on average while requiring substantially less training data. Our results suggest a new route to test-time reasoning: instead of learning how to evaluate arbitrary reasoning states, learn which states to expose so that the reasoner's existing verification signal becomes useful.