Probing isn't Causality in Molecular Models
Satya Pratik Srivastava ⋅ Rajeev Kumar Singh ⋅ Harshit Singh ⋅ Chundru Sharath Krishna
Abstract
Linear probes show whether a molecular representation linearly recovers a chemical variable. They do not show whether later computation uses that variable. We replace a concept's scalar component at an intermediate block. We measure the effect with a separately fitted readout at a later block. We compare the effect with matched random directions and architecture-matched random encoders. Uni-Mol and MolCLR give opposite results. In Uni-Mol, edits along a 3D-distortion direction predict the downstream energy-change sign for 97.6\% of held-out pairs. They also track the change magnitude with Spearman $\rho = 0.787$. These results exceed both controls at blocks 8 and 12. In MolCLR, pretraining improves dipole recovery by $0.080~R^2$. However, edits along the recovered direction achieve 0.535 sign accuracy. Random directions achieve 0.486--0.518 sign accuracy. With 1,072 pairs, the detectable shift from chance is 0.043 at 80\% power. Probe success therefore supports only the label ``linearly recoverable.'' We propose four evidence levels: not decodable, linearly recoverable, causally accessible, and pretraining-specific causal, for calibrating mechanistic claims about molecular representations.
Chat is not available.
Successful Page Load