Testing Whether Attention Recovers Finite-Difference Stencils in PDE Transformers
Abstract
We test whether attention weights recover a known finite-difference update in a transformer. Before attention is used to support scientific discovery, it should recover a known mechanism in a controlled setting. We study encoder-only transformers trained to predict the next state of heat diffusion and Lax-Friedrichs advection data generated by finite-difference schemes. The models predict accurately, but their attention weights neither match the corresponding stencils nor follow changes in the stencil coefficients. The same analysis recovers the stencil in attention-only controls, where replacing the learned weights with the exact stencil leaves prediction error effectively unchanged. In the trained transformers, fixing attention increases prediction error, while replacing it with the stencil causes a much larger increase. Fourier comparisons show condition-dependent agreement on frequencies present in the generated states, but the transformer response is closer to copying the current state at the tested higher frequencies. Attention therefore contributes to prediction without directly reproducing the finite-difference update, so its weights alone do not explain how the transformer performs the computation.