Auditing Conformal Prediction for Language Models Under Inference-Time Model Shift
Achref Doula ⋅ Max Mühlhäuser
Abstract
Conformal prediction (CP) for language models is typically evaluated under a \emph{matched} protocol, in which calibration and deployment use the same post-training inference configuration, such as quantization mode and inference temperature. In deployment, these choices often differ from those used at calibration to meet memory or throughput constraints or serving policies. We argue that this exposes a limitation in current CP evaluation protocols for language models. We formalize the underlying phenomenon as \emph{inference-time model shift}: a change in the conformal score distribution induced by inference configuration rather than by the input--label distribution, and theoretically characterize how it breaks split-conformal exchangeability even when the task distribution and trained model are unchanged. We then evaluate its consequences across $8$ language models, $4$ finite-label benchmarks, and $12$ inference configurations. We find that matched evaluation can substantially overstate deployment-time coverage, and that the effect is heterogeneous across models and datasets. We argue that inference configuration should be reported as a first-class experimental axis in CP evaluations for language-model systems.
Chat is not available.
Successful Page Load