CycleProof-Agent: Verifying Time-to-Failure Claims When Completed Failures Are the Scarce Resource
Abstract
A remaining-useful-life model revises its estimate every time new telemetry arrives, but each estimate is settled only when the asset actually fails, so verification is limited not by data or compute but by the number of completed run-to-failure cycles. Prognostic systems are nonetheless evaluated as though each prediction were an independent trial: errors are pooled over windows, and intervals cover5 one window rather than one equipment lifetime-a precision the evidence cannot support. We present CycleProof-Agent (CycleProof), a multi-LLM system that makes the number of completed failures an explicit input to what it executes and what it claims. Planner, tool-use and verifier roles rank a library of numerical predictors by their dependability on unseen cycles and seat a fixed-size committee whose claim is the member median; the trace that committee has already paid for then triages its own claims at no extra cost; and one conformal score per completed failure certifies an entire lifetime rather than a single prediction, with a feasibility condition that tells the system when to refuse a confidence level and how many further failures would restore it. On the public PHM 2018 ion-mill benchmark CycleProof leads every reported and executed method at macro RMSE 1,386 seconds using three of fourteen predictors, and turns the available failures into mode-specific confidence levels, an explicit not-verifiable status and an acquisition target; a controlled comparison shows that every construction escaping the condition substitutes an assumption whose cost we measure, and an unchanged replay on C-MAPSS confirms that committee and certificate transfer without retuning. Scarce physical evidence calls for an agentic answer rather than a purely statistical one, and we offer CycleProof as a step toward autonomous systems whose confidence grows with the evidence they have actually collected.