Detecting and Mitigating Non-Deterministic Answer Flips for Reliable LLM Inference
Abstract
Large Language Models (LLMs) now match or surpass human accuracy on many tasks. In some high-stakes settings, such as medical decision-making, however, reliability may matter even more than accuracy. In this paper, we identify a critical issue for the reliability of LLMs: asked the same question twice, the same model returns two different answers, even under greedy decoding with fixed decoding configuration. The cause lies in the arithmetic rather than in the sampling: floating- point non-associativity under dynamic batching perturbs the logits, a near-tied token decision reverses, and the divergence propagates through the reasoning that follows. The answer token at the end of that chain is therefore conditioned on different reasoning, and the answer it reports differs from the one the same input produced before. We call the resulting change of answer flipping, which is both widespread and repeatable. It appears on every model and benchmark we test, for example, multiple greedy evaluations of Gemma-4-E4B-It on MMLU-Pro, identical in inference configuration, leave 29% of the questions with more than one answer. This paper gives the first sample-level analysis of flipping and the first predictor of it. Specifically, the flipping signal can be read off from the internal consistency of reasoning. The predictor identifies which samples flip from a single run alone, on benchmarks absent from its training as well as in the domain. We then illustrate three uses of such a score: lowering the flip rate of a serving system, raising accuracy at a fixed compute budget, and estimating the run-to-run standard deviation of a benchmark score from a single evaluation, in each case ahead of the strongest alternative available at the same cost.