Data-Driven Ranking and Typology of African Agrifood Systems Using Dimensionality Reduction
Abstract
Continental food-policy monitoring is a scientific task that AI-adjacent pipelines already perform: composite indices, typologies and dashboards consume public statistics and emit rankings that steer institutions. We ask two questions of any such system: how should the underlying data be organized so that models (and increasingly agents) can use them, and how should evaluations determine which downstream conclusions are reliable? We answer with a fully audited case study at continental scale: 47,965 sourced observations for the 54 African Union member states (2000–2025), assembled under a curation protocol in which every value traces to a preserved raw API response or an identified document, candidate extractions are accepted only when they reproduce published anchor values exactly, conflicting vintages are preserved rather than averaged, and missingness is coded as an explicit, informative category. Readiness choices demonstrably determine downstream behaviour: a predecessor design with five binary indicators collapsed 47 of 54 countries into one identical profile (silhouette 0.89, an artifact of duplicated rows), whereas the re-engineered benchmark space yields 35 distinct profiles and reveals that the pattern of missing data forms its own dimension of continental accountability. The evaluation suite (bootstrap rank intervals, cluster-stability auditing with multimodality and matched-null tests, validation against outcomes withheld from construction, permutation nulls for small-n learning, and anchor year sensitivity) determines which claims survive: rankings are trustworthy only as intervals, the continent is a continuum rather than a set of types, and a 25-year regime gap shows no robust narrowing. We distill the protocol into transferable data-readiness practices for policy-facing scientific AI, and release the pipeline, archived raw data and an end-to-end verification audit.