When Is a Task Library Enough? Coverage and In-Context Forecasting
Ivan Habib
Abstract
When a temporal model forecasts an unfamiliar system, is it inferring its dynamics from context or retrieving a nearby system from pretraining? We test the retrieval explanation against the model's actual finite task library. For context $C$, $N_2=\exp D_2(Q_C\Vert P)$ is the effective number of prior tasks needed to resolve the posterior, and $\mathcal{R}=M/N_2$ is the library's achieved coverage. A preregistered AR(1) intervention confirms that raising $\mathcal{R}$ improves posterior approximation and multi-step forecasts, including against a shuffled budget-matched control. We then train a Transformer on a fixed library of coupled eight-dimensional dynamical systems. With 4,096 training systems, its longest-context forecast is within $0.1\%$ of continuous-Bayes risk even though the library is 6.84 bits short of posterior resolution; its predictions strongly favor continuous inference over exact Bayesian retrieval from the library ($S=0.973$). A separate linear-regression study finds the same qualitative separation at an 11.42-bit gap. Coverage turns ``inference versus retrieval'' into a falsifiable prediction comparison. A matched regression comparison sharpens the result: the Transformer comes within $5\%$ of Bayes risk, whereas a parameter-matched GRU remains $30\%$ above it, even though both have moved decisively away from literal lookup.
Chat is not available.
Successful Page Load