A Research Pipeline That Grades Itself: Forecast-Scored, Verifier-Gated Acceptance as a Meta-Science Object
Abstract
For AI-accelerated research the binding constraint is trustworthy acceptance, not generation throughput: whether a claim is verifiable, and whether the producer knows how likely it is to be right. The artifact defended here is a running research-production pipeline that treats its own outputs as objects to be forecast, gated, and scored. On the strongest auditable subset, 36 resolved forecasts whose seal row first appears in Git before the recorded resolution, the Brier score is 0.244 (descriptive seeded-resampling interval [0.202, 0.287]); the mean forecast is 0.504 against an outcome rate of 0.556. The broader ledger-time-ordered record has 45 resolved of 79 strict seals (Brier 0.215), with 33 open, one void, and nine resolved seal rows lacking commit-before-resolution evidence. Because resolution is outcome-selective and forecasts are temporally dependent, the intervals are descriptive rather than coverage guarantees. A typed ReDerive ledger separates independently re-derived, digit-checked, and certificate-bound results. The position is that this forecast scorecard and verification ledger, not a productivity count, are what an AI-produced research artifact should ship.