Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
Abstract
Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can substantially alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification, a task that is routinely served with compact models under reduced precision and quantization in latency- and cost-constrained deployments. To the best of our knowledge, this is the first systematic study of serving-context non-invariance in text classifiers. We train 180 models and evaluate them across 3,960 serving contexts spanning discriminative, pseudo-generative, and fully generative classifier formulations, while varying padded width, inference precision, hardware, batch size, batch composition, and other execution conditions. We show that label stability can conceal substantial score instability: changing only the batch shape changes no labels across 226,104 fp32 comparisons, yet under bf16 it redistributes up to 56.7 percentage points of predicted probability mass, with the label-change rate reaching 30.6\% among examples whose initial top-two margin is below 0.01. We trace the measured numerical batch effect to shape-dependent matrix-multiplication reductions and show that batch-invariant operations eliminate the observed score differences and label changes in the tested configurations. Taken together, our results identify and quantify the previously unreported but consequential serving conditions that must be frozen for reproducible text classification, and show that the quantization, reduced-precision, and batching choices that make compact classifiers efficient to deploy carry a measurable reproducibility cost.