Geometric Relevance Is Not Sufficient: Evaluating Focused Post-Training Supervision in Diffusion Docking
Abstract
Diffusion-based molecular docking requires learning a corrective score field across a broad corruption distribution, yet individual pose errors can differ substantially in their structural consequences. This creates a post-training conflict between preserving broad score-field coverage and concentrating supervision on corrections that appear more relevant to docking geometry. This study tests whether a fixed post-training budget is better allocated broadly or focused using DiffDock as the base docking model. Focused supervision is defined using two distinct criteria: the Cartesian consequence of torsional perturbations and the conditional consistency of torsional targets. This study compares these policies with broad continuation and with a mixed objective intended to preserve global score supervision while emphasizing selected corrections. Evaluation separates fixed score fidelity, generated-pose quality, best-of-10 sampling capacity, and confidence-selected Top-1 docking performance. The starting model achieves the lowest fixed score loss, whereas broad continuation produces the strongest aggregate Top-1 success despite worse score fidelity. Broad continuation and the starting model also attain the same aggregate best-of-10 recovery, indicating that the observed Top-1 improvement cannot be attributed solely to increased sampling capacity. Focused and mixed supervision do not produce a reproducible advantage over broad continuation under the matched five-epoch budget. These results show that geometric relevance can identify structurally consequential errors without uniquely prescribing a more effective corrective supervision policy. Under post-training, score fidelity, sampling behavior, and confidence-selected utility therefore need not move together.