When Individualized Causal Predictions Are Not Actionable: Failure Analysis, Abstention and a Runtime-Checked Narrative Agent on Ovarian-Cancer Covariates
Abstract
Individualized treatment effect models can produce a score for every patient even when the data do not support a reliable treatment direction. We study this problem using four observed clinical fields from 425 TCGA-OV ovarian cancer records. Treatment assignment and potential outcomes are simulated so the benchmark conditional average treatment effect (CATE) is known for evaluation. This is a semi-synthetic benchmark and does not estimate the clinical effect of platinum treatment. We evaluate the Causal Machine Learning Output Pipeline (CaML-OP) which links treatment effect estimation to uncertainty-based abstention and runtime-checked language-model summaries. Cross-fitting cut treatment effect prediction error by 47.4% relative to raw covariates but CaML-OP still did not outperform a simple S-Learner. Across all methods, the best accuracy for the direction of treatment effect was only 0.5988. Using the known simulated effects for conformal calibration, at a 95% coverage target, the evaluated procedure abstained on 99.4% of cases. We then tested a generator-verifier agent on 160 synthetic states. Free-text responses sometimes added an unsupported treatment direction. No unsupported direction was observed in the typed direction fields in this evaluation. Better predictions did not make individualised recommendations actionable. Abstention and runtime checks can limit unsupported claims but they do not make the underlying estimates clinically valid.