Attributing Cohen's d: Training Data Attribution for Disease-Related Effects in Normative Age Biomarkers
Abstract
A common approach to predicting disease risk is normative age modelling. Normative age models are trained on subjects assumed healthy and accurately fit their ageing trajectory. At inference, we obtain accurate predictions for healthy subjects and observe deviations for patients. We interpret the resulting gap between predicted and chronological age as a signal for disease risk. Here, we attribute the disease-related effect size of the age gap directly to individual training samples, rather than using a prediction-level loss as the attribution target. For Cohen’s d, the resulting closed-form influence functional, validated against leave-one-out retraining, ranks training samples by their effect on held-out case-control separation. Across four diseases and two biomarker modalities in UK Biobank, removing the 10% most influential training samples raises held-out disease-related effect size in every seed. Removing the top 10% more than doubles the metabolomic-age effect for type 2 diabetes and raises the brain-age effect for multiple sclerosis by roughly a third. These are substantial gains on the unchanged held-out evaluation cohorts. Random removal leaves effect size flat even at 50% removal, confirming the gain comes from which samples are removed, not how many. We show that influential subjects carry disease-specific, subclinical phenotypes. Attribution against type 2 diabetes, for example, flags elevated HbA1c, a marker of blood sugar control, not a generic marker of poor health. We release our influence-function package and our matching package for reproducibility and reuse, anonymised for review.