A low-rank decoder bottleneck bounds reliability in foundation-model perturbation prediction
Zheyu Zhang ⋅ Yuanhao Huang ⋅ Fan Feng ⋅ Jie Liu
Abstract
Foundation-model-based predictors of perturbation response are evaluated primarily by aggregate accuracy, yet lack principled per-prediction reliability criteria. We audit the state-transition (ST) decoder~\citep{adduri2025predicting} shared by recent predictors across three encoders and four cell lines under CRISPRi perturbation, and identify a rank-$1$ attention bottleneck whose leading direction maps in gene space to the same cycling/p53 stress-response program in every cell line tested. The bottleneck implies a distance-defined reliability boundary: per-prediction quality degrades smoothly with the gene-space distance from a model's predicted perturbation response ($\Delta$, the predicted change relative to control) to its nearest training $\Delta$, and biological-family identity, not training-pool size, drives held-out quality. We propose \emph{predicted-NN-distance}, this same distance computed from model output alone, as a per-prediction trust signal that tracks per-perturbation quality at Spearman $|\rho| \approx 0.81$ (oracle $0.88$) and admits a single pooled threshold that transfers across cells and encoders under leave-one-out evaluation. The signal outperforms a simpler predicted-magnitude baseline in most cases. Per-prediction reliability is bounded by a low-rank decoder geometry and by gene-space coverage of training perturbations, not by aggregate benchmark accuracy.
Chat is not available.
Successful Page Load