Litmus: Auditing a Synthetic Query Simulator for RAG Evaluation
Abstract
A synthetic query generator is a user simulator. It picks the questions a retrieval-augmented system is tested on, so every conclusion about that system rests on its choices. Those choices are rarely reported and almost never checked. We argue that a generator should be built to report on its own failures, and we use \textsc{Litmus}, a query simulator we built, to show what that buys and where it stops. Three failures can void a generated evaluation set. The questions may be answerable from memory, so they test recall and not retrieval. The judge that scores them may have an unknown error profile. The set may cover a narrower range of questions than its configuration asked for. \textsc{Litmus} measures all three from its own output, and on our corpora it fails two of them. The interesting part is how we found out. A diagnostic over the judge's own verdicts identified a redundant scoring question, and human annotation on an independent corpus identified the same one. Some failures of an instrument are visible in its output, and some need evidence from outside the loop. We say which of ours fall on each side and release code.