Passing the Benchmark, Failing the Patient: A Reproducible Validity Audit of Healthcare Agentic AI Evaluations
Abstract
Agentic AI systems are increasingly proposed for healthcare and mental-health deployment, yet this workshop's own call for papers observes an "evaluation blind spot in deployed healthcare agentic AI, where systems with strong technical metrics fail in clinical practice." We conduct a pre-registered reporting audit of validity practice in the healthcare and mental-health agentic AI literature rather than proposing a new framework. We define a reproducible sample frame over the public arXiv API (six fixed queries, retrieved 2026-08-30; 281 unique candidates, automated screening to 232 eligible, a seeded random sample of 35, and a full-text eligibility check yielding N=31 scored papers). We build a 15-item rubric by cross-walking three independently verified frameworks: the five-pillar validity taxonomy of Salaudeen et al. (2025), the benchmark lifecycle criteria of BetterBench (Reuel et al., 2024), and the Agentic Benchmark Checklist (Zhu et al., 2025). Scoring every paper in the sample with evidence-quoted extraction (author-verified on a 23% spot-check subsample), we find strict per-pillar coverage below 30% throughout: Content 26%, Criterion 19%, Construct 6%, External 6%, Consequential 3%. The single rarest practice in the rubric is testing across more than one site, population, or demographic subgroup (10% of papers). We release the query set, screening code, sampling seed, rubric, and full scoring sheet as a reusable artifact for future audits.