What Carries the Calibration Signal in Incremental Question Answering? A Pre-Specified Confirmatory Decomposition
Abstract
Incremental question answering systems read a question in stages and repeatedly decide whether to answer now or wait for more information. That decision depends on confidence. A natural idea is to estimate confidence not only from the current score but also from the answerer’s recent trajectory, such as whether the answer has persisted or its score is rising. We audit how much of the apparent benefit of those features survives a deployment-aware evaluation. We develop the analysis on a development split, freeze eight directional hypotheses, and run the protocol once on an untouched test split with 63,050 prefixes from 1,953 questions. A full model improves Brier score by 0.02253 over a linear base-feature estimator, but 29% of that gain comes from look-ahead features unavailable at decision time, 23% from changing estimator class, and 47% from causal trajectory features. Within the causal portion, score movement contributes more than answer stability on the confirmatory split. The complete causal bundle gives no detectable Brier improvement for the tested generative answerer. At the policy level, the elaborate estimator also does not reliably outperform simpler base-feature estimators. Finally, we identify an answer-identity representation bug that inflated our own generative stability estimate roughly ninefold. The results show how an apparently strong calibration result can shrink or change meaning when feature availability, estimator capacity, answerer choice, and the downstream decision are evaluated separately.