CARD: Internalizing Expert Critique into Reinforcement Learning for Deep Search
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for post-training large language models (LLMs) on complex reasoning and multi-step agentic tasks. However, the standard scalar reward signal is sparse and uninformative: it conveys whether a trajectory succeeded but not which reasoning steps failed or how to correct them, severely limiting sample efficiency on long-horizon tasks where successful trajectories are rare. We propose \textbf{CARD} (\textbf{C}ritique-\textbf{A}ugmented \textbf{R}einforcement \textbf{D}istillation), a training framework that amplifies the learning signal from a \texttt{verify_answer} tool within the GRPO objective via two complementary signals. First, a token-level self-distillation reweighting measures how much each token's generation was shaped by the expert critique: tokens that merely echo the critique text are masked, while critique-driven revision tokens are amplified. Second, a rejection-sampling distillation loss maximizes the likelihood of critique-free trajectories for responses where the critique provably improved the answer, directly accelerating the internalization of critique-guided behavior. Together, these two signals distill critique-augmented reasoning into the model's base policy without requiring additional rollouts or separate reward models. Experiments on four long-form deep search benchmarks and three multi-hop question-answering benchmarks demonstrate consistent improvements over the GRPO baseline. On the deep search benchmarks, CARD improves average performance by 40\% over the Qwen3-8B base model and by 6\% over GRPO; with Qwen3-14B, the gains are 46\% and 7\%, respectively. These results establish expert critique as an effective and scalable source of rich text supervision for reinforcement learning of long-horizon research agents. Furthermore, we show that \texttt{verify_answer} can serve as a test-time self-revision scaffold: at inference, the tool returns only task rubric criteria---without expert feedback---prompting the model to reflect on its draft and self-revise, yielding consistent performance gains over the no-VA baseline.