Differential Item Functioning as an Item-Level Diagnostic for LLM Benchmarks
Abstract
Comparative claims about large language model (LLM) performance are typically based on aggregate benchmark scores. This treats models with similar scores as comparable on the items the benchmark contains, an assumption that can fail if models from different populations succeed on systematically different items. We test this assumption using Differential Item Functioning (DIF), a psychometric diagnostic that flags items whose relative difficulty differs across groups after conditioning on overall performance. Applying Mantel--Haenszel DIF analyses to six widely-used LLM benchmarks across two widely-used open-source model families (Llama and Qwen), we find that many items behave differently across families even among models with the same total benchmark score, with some items favoring Llama-family models and others favoring Qwen-family models. The flagged items concentrate in benchmark-provided categorical content (subjects, subtasks, problem types), and choosing items based on which family they favor can substantially amplify, attenuate, or reverse aggregate family-level comparisons. Our results show that aggregate benchmark scores can summarize different item-level performance patterns across model populations, and that DIF offers a useful diagnostic for surfacing this variation.