Screening Without a Blood Test? Machine Learning for Non-Invasive Anemia Risk Flagging in Nigeria
Abstract
Maternal anemia remains a major public health burden across sub-Saharan Africa, yet its diagnosis depends on hemoglobin measurement, which many community health posts and mobile outreach programs cannot reliably perform because of the cost of hemoglobinometers, reagent requirements, and workforce training gaps. A cheaper alternative is to flag likely anemia risk using information that health workers can collect at intake, such as age, education, household wealth, and living conditions, without drawing blood. However, it remains unclear whether such indirect, non-invasive predictors contain sufficient signal to make this approach clinically useful or whether they are too weak to responsibly substitute for direct testing. This study evaluates three lightweight machine-learning classifiers, prioritizing the trustworthiness of the resulting screening approach rather than assuming that any improvement in predictive accuracy justifies replacing direct testing. Thes study used the 2018 Nigeria Demographic and Health Survey Couples Recode file, restricting the sample to 8,061 couples with a valid anemia biomarker result for the female partner. After removing incomplete records, 7,843 observations remained, while 218 were excluded because of missing anemia data. Anemia status (mild, moderate, or severe versus not anemic), based on altitude- and smoking-adjusted hemoglobin levels, was recoded as a binary outcome, with 59.3% classified as anemic. Predictors were restricted to variables that can be collected without laboratory testing: both partners' age and education, household wealth quintile, household size, urban/rural residence, current pregnancy status, and current employment status. A household environmental risk score ranging from 0 to 3 was also constructed by summing indicators for unimproved drinking water, unimproved sanitation, and solid cooking fuel use. The data were divided into 80% training and 20% test sets using stratified sampling and a fixed random seed. A depth-limited decision tree, logistic regression, and a small two-layer neural network (16 and 8 units) were evaluated using test accuracy, F1-score, model size, inference latency, and five-fold cross-validation accuracy. The neural network was additionally converted to TensorFlow Lite and quantized to int8 to assess its suitability for on-device deployment. Model interpretability was examined using decision-tree feature importance and logistic-regression coefficients. None of the three models meaningfully outperformed the 59.3% majority-class baseline. Logistic regression achieved the highest test accuracy (61.1%; F1 = 0.733), followed by the neural network (60.7%; F1 = 0.731) and decision tree (60.4%; F1 = 0.728). Five-fold cross-validation produced similarly modest performance, with mean accuracies of 60.5% ± 0.8 for logistic regression, 60.0% ± 1.2 for the decision tree, and 59.6% ± 1.0 for the neural network, indicating that the near-baseline performance was stable rather than attributable to a particular train-test split. Logistic regression produced an ROC-AUC of 0.595, sensitivity of 90.2%, and specificity of 18.6%, indicating limited discriminatory capacity and a strong tendency to classify participants as anemic. Household wealth quintile was the strongest predictor in both interpretable models, followed by women's age and education, whereas the household environmental risk score and total children ever born contributed comparatively little. Quantization reduced the neural network size from 29.0 KB to 3.74 KB without meaningful loss of accuracy. These findings do not support replacing hemoglobin testing with demographic and household proxies alone. Although the models achieved high sensitivity, their poor specificity and near-baseline discrimination indicate that socioeconomic and household characteristics lack sufficient anemia-specific physiological signal for reliable standalone screening. The findings instead support a two-stage screening strategy in which a compact, non-invasive model is used to identify individuals for subsequent point-of-care hemoglobin testing. This proof-of-concept demonstrates that model efficiency can be achieved for low-resource deployment, but predictive usefulness requires the incorporation of more direct physiological or anemia-specific indicators. Given the cross-sectional and Nigeria-specific nature of the data, further validation using prospective and clinically diverse datasets is warranted.