ImputeGuard: A Safety Harness for Agentic Missing-Data and Health-Survey Analysis
Abstract
AI agents can correctly record a scientific requirement yet violate it during execution. We present ImputeGuard, an analyst-facing safety harness for agentic missing-data and health-survey analysis. Rather than introducing a new imputation algorithm, it records when an operation is scientifically permissible and verifies that execution stays consistent with that decision; safety here means not accepting outputs that violate encoded constraints, not clinical safety or general scientific correctness. We operationalize the reasoning-execution gap (REG) as a correctly specified rule followed by an execution candidate that violates it. On the public 2023 National Survey of Children's Health, matched typed and free-form execution reveals interface-dependent REG, including 0% versus 25.53% violations in one matched setting, while an actual-value extension finds no uniformly superior interface. Unchanged-candidate replay demonstrates deterministic containment through repair, blocking, or review without model regeneration, and the live demo exposes the trace from recorded rule to conflicting execution, verifier finding, and accepted output with auditable provenance. ImputeGuard makes plan-execution discrepancies observable and enforceable within its encoded policy coverage, with no guarantee for incorrectly specified or unencoded rules.