Conformal Prediction with Paraphrase-Aware Scoring for LLM Uncertainty Quantification
Abstract
Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods often fail to remain stable under simple, meaning-preserving perturbations. We identify an important source of instability: semantically equivalent paraphrases of the same input can induce substantial variability in predictive confidence, even for methods with formal guarantees such as conformal prediction. To address this issue, we propose a paraphrase-aware UQ framework that explicitly enforces invariance to semantic rewordings. Instead of relying on a single input, our approach constructs uncertainty scores by aggregating predictions across a set of its paraphrases, using a lightweight proxy model to produce calibrated and comparable outputs. This aggregation yields uncertainty estimates that are both more stable and more informative, while remaining compatible with conformal calibration techniques to retain coverage guarantees. Across multiple UQ benchmarks and model families, our method achieves nominal coverage with smaller and more stable prediction sets under paraphrase perturbations. Code is available at https://anonymous.4open.science/r/PA_Score-8C0D.