Evaluating Pathology Foundation Models for Grading Cervical Cancer Precursor Lesions
Abstract
Evaluating Pathology Foundation Models for Grading Cervical Cancer Precursor Lesions Background: Pathology foundation models (PFMs) such as UNI, Virchow/Virchow2, Prov-GigaPath, CONCH, Phikon, etc., are visual transformers which are self-supervised, and pre-trained on hundreds of thousands to millions of whole-slide images (WSIs) image patches. PFMs have shown that they are promising replacements for task specific architectures within computational pathology. PFMs produce task-agnostic embeddings of a given image, which have proven to provide strong transferability across diagnostic and prognostic endpoints when minimally fine-tuned with few task-specific exemplars. Cervical intraepithelial neoplasia (CIN) grading is a challenging case for many transfer learning scenarios, as boundaries in grading are continuously variable rather than discrete; even expert pathologists show significant inter- and intra-observer variability when grading CIN/SIL, making them good candidates for tools that assist in more reproducible and consistent grading. Although PFMs are now abundant, league-controlled, head-to-head comparisons on cervical histopathology grading are still lacking. Methods: Eighteen publicly available PFMs were benchmarked as frozen pre-trained extractors of patch-level embeddings for CIN grading. Cross-validation of 5 patient groups was implemented on 49 patients (949 patches) to ensure that no patient’s patches appeared in both train and test. The models’ embeddings were assessed with 4 different probing methods: logistic regression (LR), k-nearest neighbors (k-NN, k=20), a linear layer, and a one-hidden-layer MLP. The models were trained and early stopping were implemented on a different patient split. The models were evaluated on their balanced accuracy (BA), macro-F1 (m-F1), and macro-Area Under the Receiver Operating Characteristic Curve (m-AUROC). Results: The strongest-performing models—virchow2, gigapath, kaikovitb16, conch, kaikovitb8, and kaikovitl14—achieved balanced accuracies between 54.2% and 56.6% when evaluated with logistic-regression probing, compared with a 25% chance level. Across the 18 models, the MLP and logistic-regression probes each produced the best result for 8 models, while k-NN and the plain linear probe were each best for only one model. The confusion patterns were also fairly consistent across the best-performing models. Grades 0 and 3 were generally well distinguished, whereas most of the errors occurred between the two intermediate grades. Macro-AUROC (one-vs-rest, four-class average) for the top 6 models ranged from 79.9% to 81.9% (virchow2: 81.9%, gigapath: 81.1%, kaikovitb16: 81.5%, conch: 81.3%, kaikovitb8: 80.7%, kaikovitl14: 79.9%), placing all evaluated models in the "good discrimination" band and considerably closer together than their balanced-accuracy spread alone would suggest. This AUROC–balanced-accuracy gap (~80% vs. ~55%) indicates that embeddings reliably rank patches by grade likelihood even when the corresponding hard 4-way classification decision is less consistent, consistent with the adjacent-grade (1↔2) confusion observed across models. Embedding dimensionality also appeared to have only a weak relationship with downstream accuracy. In particular, several encoders with fewer than 1,000 dimensions performed as well as or better than the largest model evaluated, which had a 3,072-dimensional embedding. Future Work: Future work will focus on several main tasks. First, we plan to address class imbalance more directly using approaches such as class-balanced or focal losses and resampling. We will also consider metrics that provide more detailed information about class-specific performance, including macro-averaged measures and per-class calibration statistics, rather than relying on balanced accuracy alone. These approaches may help reduce the errors involving the intermediate grades, particularly because some of these classes contain fewer samples. Second, we plan to reformulate the classification task using the clinically standard three-tier Squamous Intraepithelial Lesion (SIL) framework: Normal, Low-grade SIL, and High-grade SIL. This would consolidate the current four CIN grades into a simpler and potentially more clinically actionable target. The revised task will be evaluated using ordinal-aware measures, such as quadratic-weighted kappa, alongside balanced accuracy and AUROC. This formulation may better reflect established clinical decision thresholds while also accounting for the ordered nature of the grading categories. Third, the dataset's dual-pathologist annotations will be used to quantify inter-observer variability (Cohen's/weighted kappa, percent agreement by grade) and to test whether model errors preferentially fall on patches where the two pathologists themselves disagreed — a step toward benchmarking PFM-based grading against a clinically realistic, rather than single-rater, reference standard.