Small Gains Have Coordinates: Anchoring Strength, Ensemble Size, and Metric Choice in Adapting a Molecular Simulation Foundation Model
Harsha Poonepalle ⋅ Vedant Kalipatnapu
Abstract
The improvement reported for an adapted molecular simulation model depends on two choices that papers rarely sweep and often leave unstated: how far the weights moved from the source checkpoint, and how many stochastic samples were averaged before scoring. We vary both for one small adaptation of MarS-FM, and the second dominates: at $K=4$, the count such evaluations commonly use, roughly 83% of our gain is a $1/K$ Monte Carlo artifact of averaging few samples rather than a prediction that lands closer. The adaptation fine-tunes on mdCATH frame pairs separated by exactly 50 saved frames and keeps only 7.35% of what fine-tuning learned, which lowers endpoint error by $0.03259$ Å across all 28 eligible validation proteins, improves 27 of them, and passes a seven-observable ensemble-retention screen that every more aggressive adaptation we tried failed. But raising the samples per prediction from $K=4$ to $K=16$ leaves every protein improved while cutting the average gain by 42.4%, from $0.02401$ to $0.01383$ Å. Decomposing the error separates the artifact from the effect: what remains is $K$-independent, sitting at $0.2259$ Ų under a $1/K$ fit and corroborated by a strictly proper scoring rule. The effect is genuine and about half what the $K=4$ number advertises. Nothing in the decomposition is specific to molecular dynamics, and recovering it needs no new training. All of this rests on validation data at a single lag and temperature, with the test partition left closed.
Chat is not available.
Successful Page Load