ATOM2: An Accuracy-Cost Study of Sparse Spatial Attention in Molecular Trajectory Operators
Abstract
Molecular trajectory operators predict multiple future configurations in parallel, but dense atom–atom attention limits their application to larger systems. We introduce ATOM2, a sparse extension of ATOM that restricts each atom to at most 32 spatial neighbors within a fixed radius while retaining attention across all forecast times. On a retrospective four-fold TG80 development panel containing 31 validation molecules, an eight-layer ATOM2 model reduces two-step coordinate MSE by 1.09% relative to a six-layer dense baseline; the seven-step improvement is only 0.10%, with several horizons regressing. On a 42-atom capped-peptide development trajectory, ATOM2 reduces seven-step MSE by 6.18% but fails prespecified horizon and structural criteria. On 87-atom Stachyose, a single-seed comparison yields a 48.08% reduction in seven-step MSE, and a matched eight-layer dense control also underperforms ATOM2, although independent replication remains pending. The computational benefit is clearer: on synthetic systems evaluated on an RTX 4090, sparse attention permits forward-plus-backward execution at 2,048 atoms versus 512 for dense attention; at 1,024 atoms it is 4.24× faster and uses 9.42× less incremental peak memory. These results establish a practical scaling path for ATOM while showing that trajectory-quality gains remain system-dependent and require independent confirmation.