Quantizing Pathology Foundation Models for Local Deployment: Clinical Robustness and Backend-Dependent Efficiency
Mohammad Saber Pourheydari ⋅ Chad Vanderbilt ⋅ Christoph Lippert ⋅ Gabriele Campanella
Abstract
Local execution of pathology foundation models could reduce dependence on remote compute, but only if quantization preserves downstream clinical performance and translates into real gains on the deployment hardware. We evaluate seven pathology encoders across ten detection and biomarker tasks under FP16, BF16, FP8, INT8, and INT4 post-training quantization, replacing only the tile encoder while keeping the full-precision Gated-Attention multiple-instance learning head fixed; INT2 is evaluated separately as a retrained stress test. Across 350 frozen-head encoder–task–precision combinations, mean AUC changes range from $-0.00015$ (FP16) to $-0.00081$ (INT4), well within a prespecified $-0.005$ non-inferiority margin, whereas the 70 INT2 combinations change by $-0.03422$. Resource efficiency, however, is highly backend-dependent and does not track bitwidth alone: at batch 1, UNI-INT4 reaches $4.76\times$ FP32 throughput on Apple Core ML and $2.46\times$ on an NVIDIA H100 server reference, but $0.89\times$ on Intel OpenVINO; H-Optimus-0 at INT4 similarly reaches only $0.82\times$ on OpenVINO. Batch-8 throughput reaches up to $8.029\times$ the corresponding batch-1 throughput on H100; the largest batch-8-to-batch-1 throughput ratios are $0.971\times$ on Core ML and $1.044\times$ on OpenVINO. In the evaluated clinical setting, FP16--INT4 can preserve downstream AUC, whereas separately retrained INT2 shows a substantially less favorable accuracy tradeoff; realized deployment gains depend on the backend and hardware rather than nominal bit width alone.
Chat is not available.
Successful Page Load