Beyond Subject-Wise Splits: Auditing Whether Wearable-Health Evaluations Can Support Their Claims
Abstract
Subject-wise and leave-one-subject-out splits reduce participant leakage in wearable-health evaluation, but they do not show whether a held-out evaluation can support its full prespecified claim. Metric requirements may be infeasible, participant or subgroup support insufficient, participant-level comparisons underpowered, multiple required criteria jointly unresolved, or evaluation safeguards unverified. We introduce GateAudit-TS, a deterministic audit applied before held-out-test access. From a structured evaluation plan, it checks metric feasibility, participant and subgroup support, participant-level power, joint success across required criteria, and supplied safeguards for grouping, leakage, provenance, replay, and test access; it then distinguishes scientific redesign from process repair. We evaluate GateAudit-TS on 120 author-designed specifications fixed before expected-label access, eight public wearable-health protocols, and a separate Gaussian simulation of 900,000 studies. GateAudit-TS matched all 120 expected outputs, while a comparator omitting only the joint-success check made the same 120 binary decisions. Only 27 of 208 protocol fields were explicitly reported or exactly derivable, and no protocol supplied all 26 fields. In simulation, all six conditions approved by the audit had joint-success intervals above 0.80, whereas two of 12 conditions rejected only by the joint check also exceeded 0.80 under favorable correlation. These results support rule conformance, reporting-gap identification, and behavior under the specified model—not predictive superiority, clinical validity, or external validation.