From Document QA to Enterprise Tool Use: A Paired Outcome-and-Trajectory Study
Abstract
Enterprise-agent pass rates do not reveal whether a system found the declared sources, represented every requirement, or verified the artifact it delivered. We pair outcomes and trajectories from a base Kimi-K2.7 checkpoint and the same checkpoint after low-rank post-training on 1,465 expert-authored professional-PDF tasks. Training uses privileged-rubric distillation followed by criterion-level reinforcement learning and contains no GDPval-style file-and-tool loop. On 100 paired GDPval tasks, strict success changes from 5% to 10%, mean score from 76.4% to 78.0%, and 61 tasks improve versus 27 regress and 12 tie by normalized score (p=3.7e-4 on 88 non-ties). The same update changes GDP.pdf strict success from 11.1% to 24.2% on 99 matched tasks. Among 94 GDPval pairs with mutually gradable artifacts, pre-build source coverage increases from 76.1% to 89.9%, initial requirement coverage by 12.7 points, and executable post-write checks from 5.3% to 14.9%, despite 6.3 fewer tool calls per task. These measures localize the associated change to source discovery, requirement tracking, and intermediate verification rather than greater interaction. Because each checkpoint has one saved attempt and GDPval execution policies differ, we report descriptive cross-environment evidence rather than a causal transfer estimate. Code, prompts, and measure definitions will be released.