When the Metric Lies: Silent Evaluation Failures in RL Exploration
Abstract
A flawed evaluation metric can produce plausible numbers that support the wrong conclusion without raising any error. We document four such failures from a single project evaluating exploration in a complex game environment, all under a pre-registered protocol: a per-step rate that rewarded early death, whose correction moved a headline contrast from +33.7% to +7.6%; three coverage denominators silently mixed across runs; a noise floor high enough that one uniform-random agent beat an identical one at p = 0.033, on a comparison whose true effect was zero by construction — and whose comparator was itself the high draw of three, reversing the sign of our only negative result; and a per-turn rate that rose roughly fourfold across training while the count it divides showed no detected trend, ρ = −0.045. Two share one mechanism — a rate whose denominator counts the agent's own interaction, where episodes can end early. We give that mechanism analytically, show it is a property of the estimator rather than of our environment, and supply a short simulation in which it inverts the ranking of two agents outside any RL setting. We separate these definitional failures from the variance failures that dominate existing discussion of RL evaluation, and distill five checks.