Execution Signals for Scientific Agent Rewards: Prediction, Selection, and Transfer
Abstract
Scientific agents leave execution records that look useful as rewards; predicting a grade, selecting an answer, and guiding a budgeted decision are distinct tasks. On the original BixBench archive (7,632 graded answers, 975 pools, 45 held-out capsules), we audit whether inexpensive structural execution features add value beyond a matched reduced scorer. They change Brier loss by +0.0013 [−0.0059, 0.0079] and fixed-pool selection utility by −0.36 percentage points (pp) [−1.67, 0.93]; boosted comparisons are unresolved. A label-derived synthetic-signal check shows that the pipeline detects a clean 2.58 pp gain in 100% of replicates, with 0 of 20 false positives. In a pre-declared, scripted (not an LLM) policy paying one budget unit per revealed archived grade, execution features change the logistic scorer’s archived-positive yield at β=0.5 by −0.29 pp [−1.49, 0.91]: practically equivalent to zero within the original protocol’s ±2 pp relevance margin (minimum detectable effect 1.89 pp), so the pre-declared falsifier fired. The comparator is weak: full logistic does no better than random allocation, and answer consensus reaches 3.52 pp higher yield (descriptive). Post hoc, capsule-level score and grade means correlate (Spearman 0.67), but the within-capsule AUROC, 0.47 [0.37, 0.57], is imprecise and includes chance. Multiple choice without refusal changes the target: an open-trained scorer has 0.0450 higher Brier loss than a protocol-trained counterpart. Conclusions are bounded by one reused archive whose questions were later revised, 327 attempts without grades, inherited LLM grades rather than validity labels, and no semantic-judge baseline. We release code, data, and decision records.