More Compute, Worse Decisions: When Best-of-K Reward Estimation Breaks Constraint Certification, and How to Diagnose It in Advance
YUFAN XU ⋅ Yujia Zhang
Abstract
Molecular reinforcement learning pipelines usually score a candidate structure by generating $K$ conformers or docking poses and keeping the best score. This best-of-$K$ reduction is a biased extreme-value estimator of affinity, and its bias grows with the variance of the candidate's own evaluations: $\mathbb{E}[\max_{j\le K} s_{ij}] \approx \mu_i + a_K\sigma_i$ ($R^2 \ge 0.97$ for $K \le 16$, falling to $0.94$ at $K=32$). The same relation holds for two different oracles and four docking banks covering three targets, two protein folds, and two empirical scoring potentials ($10{,}464$ replicate evaluations). Because the coefficient $a_K$ grows with the sample budget, spending more compute per candidate pushes the operational reward further from the quantity it is meant to estimate. Ranking benchmarks hide most of this ($\rho \approx 0.96$), but threshold-based constraint certification does not: the rate at which infeasible candidates are wrongly certified rises from $6.7\%$ at $K=1$ to $25.6\%$ at $K=16$ when measured against held-out replicate means. Recalibrating the threshold so that every configuration keeps $80\%$ detection power does not repair this ($2.9\% \rightarrow 6.1\%$), so the failure reflects lost information rather than a calibration offset. Its size depends on whether the constraint refers to a central or an extremal property, and on the noise-affinity correlation $\rho_{\sigma\mu} = \mathrm{corr}(\sigma_i, \mu_i)$. The latter can be estimated from replicate runs alone, with no ground-truth labels, and its sign predicts the direction of the compute effect: it is negative in the three banks where compute degrades certification and positive ($+0.54$) in the bank where certification improves. A 144-run permutation experiment that holds the marginal distributions fixed shows the relationship is causal. Against experimental ChEMBL IC$_{50}$ values, more compute weakens the correlation between best-of-$K$ and measured potency while strengthening that of the sample mean (pooled contrast $+0.034$, $p = 0.02$). Of eleven reduction operators, the sample maximum is never the best choice once $K \ge 4$.
Chat is not available.
Successful Page Load