Reward Observability under Changing Biomedical Analysis Contracts
Abstract
A reward for biomedical analysis must distinguish a convincing result from an analysis that answers the stated scientific question. We audit this with 1,440 executed workflows on four public biomedical datasets and 14,400 dependent execution–contract records. A contextual random-forest reward reaches aggregate AUROC 0.924–0.958 yet makes two opposite transfer errors: it rejects all 200 valid records of a held-out composition, and on the repeated-subject dataset it accepts subject-leaking validation, a constraint absent from reward training. Post hoc, on the primary contract alone, it accepts 14.2–31.7% of invalid workflows on three datasets, against at most 4.875% in the all-contract mixture. A scripted generative agent that executes typed workflow choices (model, preprocessing, validation, endpoint) and sees only the learned reward submits no invalid workflow at 32 executions on the three datasets where the reward had training support, but leaks subjects in every Parkinson split. In a pre-declared test, a local open-weights language-model agent with the same actions, budget and reward, plus the contract in its prompt, leaked in 0/30 Parkinson runs (95% CI [0.000, 0.116]) against 28/30 for the scripted agent: it followed the stated contract rather than the reward, and explored little. The amplification concerns reward-driven optimizers that do not read the contract. Composite support removes rejection on three datasets, within-cohort subject support only partially changes selection, and three added learners tie on both failures. This is a controlled reward audit, not a general agent evaluation. Rewards should be audited by contract, composition, training support and search budget, not aggregate discrimination.