Auditing Financial ML Backtests with Evaluation Contracts and Executable Checks
Abstract
A reported backtest metric reflects both its target generator and its evaluation protocol. We study historical simulated performance under declared information timing, execution, residual-cash return, modeled cost, comparator, and engine semantics. In the reference implementation, a version-1 contract and a canonical tape of timestamped target weights connect selected declarations to executable checks and content-hashed artifacts. A retrospective audit of a six-ETF rotation strategy reports net CAGR of 2.05% and Sharpe of -0.04 relative to the SHV cash proxy. An exact daily arithmetic identity attributes -3.05% per year to active risky allocation. A repository-frozen expansion evaluates one walk-forward gradient-boosted target generator, using five fixed seeds, and three rule-based generators over 2,011 sessions in U.S. sector and country-equity ETF panels. Retroactive signal-close assignment, which uses unavailable information, has annual-return intervals below zero in 3 of 8 family-panel cells and overlaps zero in 5. Removing an illustrative 13-basis-point charge on gross two-sided executed notional raises annualized arithmetic return by 0.57 to 1.81 percentage points across these cells. These path-conditional finite differences do not identify causal effects, prediction accuracy, future alpha, investment skill, or deployment safety. Neither panel is a temporal holdout. Empirical reproduction requires matching nonpublic source matrices, and all intervals are pointwise without a multiplicity adjustment. The workflow combines contracts, checks, contrasts, and provenance; it introduces no new estimator, theorem, or general benchmark.