Do Scalar Summaries of Foundation-Model Embedding Geometry Improve Volatility Forecasting? A Calibrated, Leakage-Aware Audit
Abstract
Frozen foundation models now supply plug-in features for temporal prediction, usually on in-sample evidence. We audit one such feature end to end, over fifteen years of US stock volatility. The feature is the geometry of each stock's recent news-headline embeddings: how far apart they sit in embedding space. Every step of the estimation is leakage-aware, except one disclosed selection step. News is usable only once tradable, evaluation is out of time with a rolling-origin check, every configuration tried enters a trial log that raises the bar, thirty stocks are held out of design, and a final year is opened once. We also calibrate the audit. A known-real persistence signal passes, every planted noise feature is refused, and planted true signals of graded strength give it a measured detection threshold. In-sample the geometry looks predictive under six of seven encoders; under the audit none is certified. The in-sample number that drew attention is matched by pure-noise features on the same data, and an effect as small as the leading candidate's would be certified less than half the time. Scalar summaries of headline-embedding geometry, entered linearly, do not survive a calibrated, leakage-aware audit of stock volatility. The same cascade, run on Amazon product reviews with one forced change, certifies a real effect through a holdout it had never opened.