A Self-Improving Forecasting Agent That Chooses Its Own Validation Window: Design Rules from a Two-Stage Pre-Registered Study
Abstract
We implement a self-improving forecasting agent: over 5 rounds it proposes edits to its own forecasting pipeline, scores them, and promotes a champion. Nothing forces such a loop to score its edits on a window fixed in advance, so we build two agents differing in exactly that. We pre-registered, before any data existed, that the self-selecting agent reports gains at least 3 pp above its sealed-test gain on ≥ 2 of 3 real streams. That claim is REFUTED (0 of 3). A second pre-registration, written before the extension data existed, promoted the original secondary outcome—does self-selection damage real generalisation?—to confirmatory primary on 7 fresh streams × 30 seeds. Its rule returns SUPPORTED (4 of 7), but the inconsistency is the finding: electricity loses −4.7705 pp and traffic −3.4780, while ETTh2 (+0.5532) and solar (+0.6792) significantly improve, and the pre-registered rival mechanism—stream family—is refuted. Adversarial review then showed A-vs-B confounded which window with how many, so we completed the 2 × 2 under a rule fixed in both directions: its verdict is MIXED-UNRESOLVED, with pooled aggregation −0.0198 pp indistinguishable from zero against adaptivity −0.9321, and the 95% prediction interval for a new stream spans zero. Swapping the pipeline for a pretrained foundation model (Chronos-Bolt) REPLICATES the pattern on 2 of 4 streams—but 2 of those 4 are in its pretraining corpus, so a corpus-clean replication remains open. An LLM-proposer arm was abandoned mid-run for order-correlated contamination, so cross-proposer robustness remains open.