Structural Blindness in Latent Data Assimilation: Representation Geometry Misleads Sensor Design
Abstract
A learned latent space supports reliable scientific decisions only if decisions made in that space remain valid in the original physical variables. We study latent Kalman-type data assimilation, where an encoder--decoder pair is trained without knowledge of the sensor used to collect observations. Can representation-level diagnostics predict which sensors give accurate physical-state filtering? In general, no. We show that any such diagnostic (a smooth functional of the averaged decoder Fisher matrix, including optimal-design criteria, posterior Cram\'er--Rao surrogates, and effective-observability scores) sees the sensor only through a \emph{Fisher--Gram coarsening} of the physical sensor Gramian, leaving a large blind spot. The failure is unconditional, an algebraic consequence of averaging the decoder Jacobian over the latent invariant measure. The positive side is conditional: when the proportionality between physical and latent filtering error is stable across sensors, latent filtering error ranks sensors accurately, with an explicit Pearson bound. Decoder smoothness provides one route to this stability, and a linear-Gaussian Riccati surrogate (built entirely from the trained latent model and decoder) empirically tracks the same ranking when full-state truth is unavailable. Experiments on five chaotic systems support both sides. Representation geometry is useful but structurally limited, and reliable sensor ranking in latent data assimilation requires filter-aware validation rather than decoder-diagnostic scoring.