Consistent Reports, Invalid Experiments: Evidence Boundaries of ML Result Validators
Abstract
A machine-learning report can be internally consistent and exactly reproducible while violating its declared experiment. We study the execution evidence needed to distinguish these cases. Among 360 constructed violations on three datasets and two estimator families, 63 expose the same canonical report as a valid counterpart, defeating even a correct-output reference. Report-local consistency detects 30 faults; lineage and independent checkpoint replay together detect 360 with no alarms on 150 valid variants, after a disclosed numerical repair. All checkers still accept 30 inadequate-specification probes. Transfer to documentation programs reveals the converse failure: positional operand conventions reject valid symmetric metrics. We introduce evidence-consistent operand binding and token alignment, coverage, and reduction checks. On 240 new tabular executions, binding detects 120 faults without false alarms, versus 40 false alarms for the positional rules. A separate loss-evaluation study uses two pretrained language models, four books, and 72 fixtures. Strict reconstruction rejects all valid float32 cases; a disclosed float64 loss-precision contrast detects 504/504 faults with 0/360 false alarms and abstains on 144 unsupported probes. Component ablations expose distinct missed faults. The new checks match a privileged correct-score oracle on these cohorts without requiring its expected score. These dependent, constructed cases support bounded conformance checks, not real-world defect prevalence, improved model quality, or a general research-code verifier.