A quarter of ChEMBL’s potency labels are bounds — and deleting them costs more than keeping them
Abstract
ChEMBL records whether a potency measurement was a determination or a bound: standardrelation is = for an exact value, > when the assay never reached a response, < when it saturated. Benchmark builds routinely read standardvalue and discard standardrelation, so “IC50 > 40,000 nM” enters training as “IC50 = 40,000 nM”. We quantify what that costs. Over 5013576 potency records in ChEMBL 36, 24.34% are censored, rising to 54.57% for Kd. Censored records sit at the edges of the tested range; median potency shifts -1.6982 log units from exact to right-censored IC50 records — partly by construction, but at a magnitude a pipeline inherits as if it were signal. We then show the field’s tooling handles this inconsistently: pchemblvalue, the convenience field ChEMBL defines only for exact relations, is populated for 96.53% of exact rows and 0.0% of right-censored and 0.0% of left-censored ones. A pipeline keyed on that field therefore drops every censored record; one keyed on standardvalue keeps them and, unless it also reads standardrelation, inherits the bias. Which of the two datasets you built is not visible in the resulting file. The modelling consequence contradicts two of our three pre-registered predictions and returns the third as a null. Across 158 single-protein targets (every one meeting a pre-registered inclusion rule), treating bounds as exact values did not cost accuracy against exact-only held-out labels, it saved it: naive minus exact-only RMSE averages -0.0777 log units. A size-matched control shows why: at equal training size the censored labels do not help (0.0465 log units against them on 149 of 158 targets), below our pre-registered floor and so reported as null. Keeping them wins on volume alone, and the pre-registered correlation flips sign under size-matching (ρ = −0.567 unmatched, 0.429 size-matched, post hoc). In this linear IC50 benchmark the Tobit correction, which exists to model those bounds correctly, instead costs 0.1077 log units against naive on identical rows, worse on 156 targets. The decomposition separates a real volume gain from a label-quality difference too small for our floor to call.