Same Criterion, Same Cohort: Paraphrase Robustness in Agentic Clinical Cohort Discovery
Abstract
Cohort discovery — turning a clinician's description of a patient population into a concrete patient set — is named as an agentic decision-support task by this workshop's call. Text-to-SQL research shows single-shot query systems lose substantial accuracy under meaning-preserving paraphrase. We ask whether that brittleness survives when an LLM agent instead plans and executes a multi-step tool sequence (filter, deduplicate, count, sample, summarize) over structured chest X-ray metadata. Across 5 cohort definitions, 5 equivalence-verified paraphrases each, and 3 trials per paraphrase (75 trials), we find no paraphrase-driven drift: every trial matches a deterministic, disclosed gold pipeline exactly (Jaccard = 1.00, N-match = 1.00, max standardized mean difference = 0.00, 0% deduplication failure). A no-tool floor confirms the agent cannot fabricate a cohort without tool access, and the null holds under harder criteria and with the deduplication tie-break unstated (30 trials each). Only a cross-model check breaks it: re-running all 75 trials on Claude Haiku 4.5 surfaces one failure, a seed that narrowed an OR query into an AND-like no-op and returned 499 patients instead of 735 (Jaccard = 0.68) with age and sex still balanced — named by our harness from the log alone. Two further findings concern the instrument itself: of 22 single-step errors injected into the gold pipeline, every one is either flagged or provably immaterial; and in 10/75 trials the agent's self-report listed tool calls its immutable log does not contain, a caution for evaluations that score agents from their own summaries.