OWCE: Revealing Language Model Complementarity via Online Weakness-Conditioned Evaluation
Hankyul Baek ⋅ Jaewon Noh ⋅ Sang Seo ⋅ Sungpil Shin ⋅ Yongsu Kim
Abstract
While modern applications of large language models (LLMs) rarely rely on a single model, evaluation still collapses each model into a single benchmark score. This simplification obscures each model's specific weaknesses and how those weaknesses are shared or unique across models. To address this, this paper presents Online Weakness-Conditioned Evaluation (OWCE), an evaluation framework that diagnoses each model's recurring weaknesses, generates weakness-conditioned follow-up items at increasing difficulty, and cross-evaluates the resulting items on all models. Beyond a scalar score, OWCE produces per-model weakness profiles, enabling direct comparison of where models fail and which others recover those weaknesses. These profiles expose pairwise structure that scalar scores leave invisible. This paper finds that each model has a best-matched counterpart that reduces its failure rate by up to $25$ percentage points. This counterpart is determined by weakness profiles rather than overall benchmark ranking, indicating that the highest-scoring model is not always the most effective partner for a given target. This paper evaluates OWCE empirically across seven benchmarks and seven models, and conducts a mathematical analysis showing that the generated items can provide more discriminative information under benchmark saturation. Code is available at https://anonymous.4open.science/r/OWCE-57E7
Chat is not available.
Successful Page Load