From Foundation Models to Tiny Programs: LLM-Discovered Time-Series Anomaly Detectors
Abstract
Time-series anomaly detection trades off predictive accuracy, computational efficiency, and interpretability, a tension that becomes acute under deployment constraints. We use a large language model not as the deployed detector but as the \emph{author} of one: an autonomous research loop repeatedly edits a short NumPy program under a leakage-free synthetic-validation objective and retains the best-scoring detector it finds. The loop discovers compact univariate and multivariate detectors that describe short windows by local spectral features and compare them with the training-region distribution through covariance-aware novelty scores. On TSB-AD these detectors take the most dataset-family first-place finishes on each reported metric, ahead of strong classical, deep, and foundation-model baselines including Time-RCD, while training no network and requiring no GPU at inference. The multivariate detector is also the fastest method in our deployment-time comparison. This separates expensive general-purpose reasoning at \emph{discovery time} from a small, auditable artifact at \emph{deployment time}, suggesting program discovery as a route from foundation-model capabilities to resource-constrained time-series inference.