When the Leaderboard Winner Is Not the Most Grounded: Uncertainty-Aware, Multi-Criteria Evaluation of a Documentation RAG Assistant
Abstract
Evaluation results are increasingly used to pick which model configuration to deploy, yet the evidence behind such choices is often a single metric, reported without uncertainty, under a single evaluation protocol. We treat this practice as the object of study in a controlled case study: a documentation-grounded retrieval-augmented generation (RAG) assistant over the Kubernetes documentation, with 20 LoRA configurations of two generators (3B and 8B) evaluated on a manually verified benchmark of 5,144 QA pairs. We measure token-level F1, LLM-judged groundedness and correctness, latency, and memory, attaching bootstrap 95% confidence intervals to every point estimate and paired-bootstrap intervals to every comparison, and we re-run the entire grid under 10 retrieval/prompting protocol variants. Three evaluation-methodology findings emerge. First, most point-estimate rankings dissolve under paired bootstrap: of six natural "winner vs. runner-up" comparisons, only three are statistically supported. Second, token-F1 and judge-based groundedness measure different constructs: the F1-optimal and groundedness-optimal configurations never coincide in any of the 10 protocol variants. Third, conclusions at the level of configuration families survive protocol perturbation even when the per-protocol optimum moves. We also show how a param-matched control separates a structural effect from a parameter-count confound, and discuss judge blinding and same-family bias. We distil the case study into practical recommendations for trustworthy evaluation of RAG systems. The benchmark and evaluation artifacts are available at https://github.com/EugPal/rag-lora-tradeoffs.