What Counts as a Good Match? A bounded, three-axis evaluation for preclinical model selection
Abstract
Choosing a preclinical model means deciding which properties of a patient’s tumour an experiment must preserve. We turn that decision into an evaluation protocol for computational methods that place patient tumours and preclinical models in a shared expression space, from PCA to transcriptome foundation models. For each patient tumour, the nearest cell lines or xenografts in that space are retrieved and scored on three axes that rest on distinct evidence: graded disease-ontology similarity, driver-mutation overlap, and donor identity. Each score is read against two references computed on the same pool of candidate samples, the random expectation and a candidate-dependent oracle, so that a low score can be traced to the method or to a database with few suitable candidates. Illustrative comparisons across primary tumours, metastases, cell lines, and patient-derived xenografts show that agreement at the disease level does not imply genomic or individual correspondence. The contribution is a protocol linking biological relevance, eligibility, training exposure, and score interpretation, not a new similarity metric.