TReDS: Trajectory-Grounded Requirement-Capability Modeling for Training Distribution Shaping in Tool-Interactive Tasks
Han Liu
Abstract
Verifiable tool-interactive agents produce rich trajectories that contain far more information than final success or failure. However, standard RL post-training with verifiable rewards often collapses these trajectories into scalar outcomes and trains over a largely fixed data stream, leaving open the question of which examples are most useful for improving the current policy. We introduce TReDS, a trajectory-grounded framework for capability-conditioned training distribution shaping. TReDS estimates multidimensional task requirements from offline probing trajectories and controlled tool-view evidence, maintains the policy capability state in the same semantic space, and uses reliability-gated online correction to form a policy-conditioned requirement representation. A scheduler then converts requirement--capability mismatch, recent utility, and stabilizing pressure into a time-varying training distribution for GRPO, without modifying the GRPO objective. We evaluate TReDS with a progressive stress-test suite spanning $\tau^2$-Bench official evaluation, $\tau$-Bench native-bridge evaluation, BFCL V3, and API Bank, using only the $\tau^2$-Bench training split for policy optimization. TReDS improves over the GRPO baseline under the same training data and backend, and shows stronger behavior retention under same-family and cross-protocol evaluations. These results suggest that RL post-training for tool-interactive agents should optimize not only how trajectories are rewarded, but also where trajectory-revealed requirements allocate training probability mass.
Chat is not available.
Successful Page Load