A Transformer-Derived Iterative Preconditioner
Patrick Lutz ⋅ Themistoklis Haris ⋅ Aditya Gangrade ⋅ Venkatesh Saligrama
Abstract
We show that a softmax-attention-only transformer trained on in-context linear regression discovers an iterative symmetric preconditioning primitive. The trained model is used only for algorithm discovery: the final method, RAPID, is an explicit preconditioner and uses no learned component at solve time. To extract this mechanism, we apply a layerwise symmetry intervention strategy which reduces the trained layers to a two-scalar update. The current Gram geometry determines the attention scores, and the resulting row update evolves the Gram matrix by congruence, reducing its condition number. We turn this primitive into RAPID, an anytime preconditioner for explicit SPD systems $Ax=b$: every prefix constructs a valid factor $P$ such that $PAP^\top$ can be used to improve the runtime of iterative solvers such as conjugate gradients. RAPID uses fast randomized Walsh--Hadamard mixing and sparse row-min corrections, giving $\widetilde O(d^2)$-time dense iterations. For a Haar-mixing idealization, we prove that RAPID reaches a target condition number in $\widetilde O(d\log\kappa(A_0))$ expected iterations. Empirically, RAPID is fast and versatile: it reaches useful condition numbers quickly, accelerates end-to-end CG solves, and improves conditioning across all tested spectral families and real-world SPD problems.
Chat is not available.
Successful Page Load