Evaluating Professional-Document Post-Training on Tool-Using Tasks
Abstract
Post-training gains are often measured in the usage context that produced the training signal, whereas deployment may jointly change tasks, observations, actions, horizon, outputs, grading, and harness policy. We present a cohort-explicit evaluation of one such compound change. Kimi-K2.7 is post-trained on 1,465 one-shot professional-document tasks using privileged-context distillation followed by criterion-reward reinforcement learning. On the same 99 source-evaluation tasks, strict pass changes as 11.1% → 11.1% → 24.2% and mean score as 66.66% → 67.75% → 74.84%, localizing the large saved-checkpoint difference to the interval after distillation. Under GDPval's separate file-and-tool context, strict pass changes from 5% to 10% and mean score from 76.4% to 78.0% across 100 paired tasks; 61 scores increase, 27 decrease, and 12 tie, and 69.3% of non-ties favor the final configuration (exact 95% CI, 58.6–78.7%). On 94 pairs with mutually gradable artifacts, differences concentrate in pre-construction source coverage, requirement representation, and executed post-write checks; arbitrary completion of the six excluded pairs cannot reverse the binary post-write-check contrast, while qualitative review finds no clear terminal-inspection contrast. We state which contrasts each comparison identifies: within-environment pairing identifies a configuration contrast; matched execution policy would identify a checkpoint contrast, which holds for GDP.pdf but not for the saved GDPval configurations; and curriculum randomization would be required to attribute transfer to the structural obligations the training tasks exercise. We propose the resulting cohort-explicit reporting protocol—separating all-task outcomes, conditional process measures, and source-checked cases—as a minimal discipline for evaluating post-training under a changed usage context.