Same Benchmark, Different Conclusions? Auditing the Stability of Kazakh Large Language Model Evaluation
Akylbek Maxutov ⋅ Nartay Aikyn
Abstract
Large language model (LLM) benchmarks typically report a single score from one evaluation configuration, assuming that minor, reasonable evaluation choices do not significantly affect results. We test this assumption for Kazakh multiple-choice evaluation by assessing 18 model configurations across three benchmarks and 11 conditions, generating 1,007,664 item-level outputs. We vary irrelevant context, instruction language, option order, and question wording, and also measure repeated-baseline disagreement, cyclic option rotations, and no-question leakage. Our findings reveal a two-level pattern: individual predictions are unstable, but aggregate conclusions remain stable. Identical baseline reruns disagree on 18.0\% of predictions on average, and validity-preserving perturbations flip 24.2\%. However, these perturbations change accuracy by only $-1.17$ points on average. Kendall's $\tau_b$ stays at least 0.935, the top configuration remains unchanged, and at most 1.4\% of baseline-separable model pairs reverse. Instability is concentrated in low-margin items: predictions with full baseline agreement flip 9.5\%, compared to 76.6\% when all three baseline runs disagree. We also find that generated "relevant distractor" context leaks answer signals and cannot support a distractor-robustness claim. Our results suggest that evaluation stability should be measured at both the item and leaderboard levels, and that perturbation-induced flips should be interpreted relative to an empirical self-flip baseline.
Chat is not available.
Successful Page Load