An Evaluation Environment Scored Against the Exact Conditional Law, with a Measured Reference Learner
Pin-Hao Chen
Abstract
Probabilistic forecasts are almost always scored against a realised trajectory, which conflates model error with sampling noise and leaves a second question unanswerable: whether a small score means the model is good or merely that the task is easy. We release a temporal evaluation environment in which the predicted object is scored against the exactly computed conditional law rather than against draws from it, and which ships a measured reference error — what a reference learner actually attains at every rung of a difficulty ladder, so a shortfall is measured against something achieved. The environment is a finite discrete limit-order market with rational transition probabilities; the target comes from exact filtering and exhaustive propagation, verified by two independent exact implementations and a sampling one. Using it we find the evaluated amendment subset solvable by a frontier model given enough test-time compute: one model's error collapses from $2.2\times$ an announcement-blind reference to below the reference learner's between roughly $3\,000$ and $9\,000$ output tokens, matching $p^\ast$ to $10^{-6}$ cell by cell on $79$ of $92$ high-effort instances, with a non-zero tail on the rest. Within a common output-token band the ordering persists and error declines with observed reasoning length at sharply different rates — an observation, not a controlled comparison, since models do not spend equally within a band. Matching that comparison's $0.010$-nat per-instance standard error from realised draws would take about $3\,300$ per history, in the least demanding of $6$ cells.
Chat is not available.
Successful Page Load