M4Bench: Evaluating Procedural Specification for Clinical EHR Derivation Agents
Abstract
Clinical research over electronic health records depends on derived variables whose definitions sit outside the schema, in item codes, time windows, missingness rules, and output-grain choices that change cohorts and downstream conclusions. We introduce M4Bench, 28 clinical derivation tasks across 15 task families on MIMIC-IV and eICU, with controls that decompose the value of procedural context for AI coding agents into specification completeness, delivery format, and schema familiarity. A Docker-isolated Codex campaign isolates specification completeness as the primary driver: an operational spec over the schema-only baseline yields +0.162 [0.072, 0.260] on eight sentinel tasks, while clinician-reviewed skills add only +0.024 [-0.034, 0.099] on top, and the combined targeted-skill effect over no-skill is +0.136 [0.085, 0.191] across all 28 tasks (26/28 improving). Only 87/760 primary runs pass all strict diagnostics, with failures dominated by output-grain and temporal-boundary errors, and key F1 moves only +0.003 task-balanced, so the gain is partial reference agreement on matched keys, not autonomous task completion or row-grain repair. Schema restructuring preserves 0.704 of the with-skill gain on gpt-5.4-mini over the 26 MIMIC-IV tasks, bounding schema-name memorization at ≤ 29.6% of the effect; decoy-skill and raw-SQL controls further rule out generic-context and reference-code-structure-only explanations. Skills are released as a versioned, clinician-reviewable markdown delivery format, one packaging among several, with task-level reference-validity audit and per-run egress logs. M4Bench releases tasks, skills, evaluation code, and per-run audit logs as a fully open-source benchmark object designed for extension and reuse.