Unreliability or Disagreement? What an Interpercentile Decomposition Can and Cannot Show
Abstract
A widely cited result reports that language models "get lost" in multi-turn conversation, and decomposes the degradation into a small loss of aptitude and a large increase in unreliability, where aptitude is a high percentile of an instruction's score distribution and unreliability is an interpercentile range. We show that this decomposition cannot support the interpretation usually placed on it. The underlying statistics are not new. Dispersion of a binary variable peaks at intermediate difficulty, a fact long known in classifier-ensemble diversity and item-response theory. They have not, however, been applied to this metric, and their consequences here are severe. Analytically: the two components are nested, since the range is the difference of the two percentiles that define aptitude and the lower tail, so they are one reparameterisation of a percentile pair rather than independent axes; at the original's own ten runs per instruction and under a binary scorer, the range saturates at its maximum for seven of eleven possible outcomes and vanishes at both extremes, making it a near-binary indicator of whether an instruction is contested rather than a graded measure of reliability, so a model that always fails is scored as reliably as one that always succeeds; and because a range is non-negative, the comparison degenerates as a baseline saturates: the reported change is forced non-negative in the limit, and its relative form is inflated well before that. Empirically, on 5,765 conversations we generated across seven models from six developers and three tasks, three of sixteen cells have a baseline unreliability of exactly zero, so their headline increase is guaranteed by construction; a null containing no reliability parameter at all overproduces the original's headline unreliability rise while underproducing its aptitude drop, inverting the intended reading; and a binomial null that sees only each cell's accuracies and run count, and nothing about the task, recovers the ordering of the per-task results (Spearman rho = 0.66). Replacing the statistic with an operational retry test separates what the decomposition conflates: on mathematics an oracle retry recovers 92% of the degradation, while on function calling it recovers 48% and leaves a gap strictly positive in seven of seven cells, with a fifth of instructions failing deterministically. We conclude that interpercentile decompositions should be reported alongside the percentiles they are built from, and that benchmark saturation must be checked before such a decomposition is interpreted.