Pretraining Buys Short-Horizon Direction: A Randomly-Initialized Control for Frozen Temporal Representations
Abstract
Self-supervised pretraining is widely reported to produce temporal market representations that predict volatility and risk. We test what it actually contributes, using a control this literature omits: an encoder of identical architecture whose weights are never trained. We pretrain a multimodal joint-embedding predictive encoder on hourly U.S. equity data and freeze it, then score it and its untrained twin on FinState-1H, a benchmark of twelve pre-registered tasks under a rolling support/query protocol, against nine tuned references including gradient-boosted trees and HAR-RV. On second-moment and tail tasks, the random encoder does as well as the trained one almost everywhere: at feature parity we can exclude second-moment gaps above roughly 0.04 and no more, which leaves the risk structure usually credited to pretraining indistinguishable from what the architecture and its inputs already supply. What pretraining buys is short-horizon direction, peaking at +0.060 over the untrained encoder at 2 hours ,the pretraining objective carries a signed-return channel; so this isolates what optimizing that objective bought not self-supervision in general. The two effects at h=1 and h=2 replicate across fold counts, reference pools, bootstrap constructions, input regimes, and two independent random-encoder initializations, and in its pre-registered run the 64-input contrast encoder beats a tuned XGBoost at one hour under that specific fold configuration, the only result in the suite to pass both controls, though this margin closes to -0.0019 at feature parity. We evaluate one encoder lineage on one market and frequency, and we report all twelve tasks, nulls included, making no trading claim: the benchmark and its controls are the contribution.