Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
Abstract
We present a benchmark for predicting exact final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction to longer executions and intermediate state, pairing shorter- and longer-trace final-output tasks with checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their reasoning-off counterparts by 33.1 to 55.2 percentage points, yet substantial gaps remain across task conditions. The strongest setting scores 93.0\% on shorter-trace final output, 77.0\% on longer-trace final output, and 65.5\% and 63.5\% on the two state tasks. Across 2,397 matched Python comparisons with identical source, changing to the longer-trace input yields 528 correct-to-wrong changes and 147 reversals. Exploratory pooled analyses associate cumulative state load with error, with weaker support within individual task conditions. Cyclomatic complexity shows no consistent positive association with error in the long-final condition. The benchmark exposes errors on longer executions and checkpoint states that short-output scores alone would leave unseen, providing related tasks and execution-generated oracles for evaluating code reasoning.