DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining
Abstract
Large language models can encounter future outcomes during pretraining, compromising historical forecasting evaluations. We introduce DatedGPT, twelve 1.3B-parameter models trained from scratch on approximately 100B tokens each with annual data cutoffs from 2013 to 2024, and DatedInstruct, instruction data grounded in each year's documents. The models retain competitive language capabilities, while headline perplexity patterns are consistent with their temporal cutoffs. On 61,290 firm-day news headlines, the lookahead-free instruction-tuned series achieves an annualised Sharpe ratio of 3.20. A signal decomposition estimates a lookahead premium of 26.4 basis points per standard deviation (t=10.65). These models support direct audits of temporal leakage in financial forecasting.