Pricing the Feedback Channel: Random Audits and Adaptive Specialist Selection in Budgeted Protein Discovery
Abstract
A discovery controller that can call several biological specialist predictors must decide which one to trust, and whether to spend part of its experimental budget on feedback that would let it decide better. We turn ProteinGym into a budgeted sequential-decision benchmark and price that choice. A controller sees five zero-shot specialists, selects 8 variants per round from a pool of 1024, and receives previously measured deep-mutational-scanning outcomes; the total budget is 56 measurements on each of 30 protein-disjoint held-out assays, with 20 paired seeds and a protocol frozen before any evaluation result was computed. Our primary contrast asks whether restricting the specialist-weighting estimator to uniformly random audit outcomes beats using all observed outcomes at the same audit allocation. It does not: the paired effect is -0.0002 normalised AUC (95% CI [-0.0010, +0.0007], p = 0.7051), a null. Reserving audit slots at all is actively costly (-0.0032, CI [-0.0043, -0.0021], 5.1% relative), and adaptive reweighting does not beat a fixed equal-weight rank ensemble, which is the strongest deployable method we tested (0.0655 vs. 0.0354 for random selection, +85%). A fixed-history diagnostic explains the null: on identical observation histories the audit-only and all-observed estimators identify the retrospectively best specialist about equally often (0.397 vs. 0.392 at round 5), both far from reliable. At this budget the bottleneck is the weak informativeness of a few dozen outcomes, not the feedback channel they arrive through, so paying for cleaner feedback buys nothing.