Proxy Metrics Can Hide Downstream Collapse: A Case Study on Compressed Cellular Segmentation
Abstract
Post-training compression of vision models is commonly screened using output-level proxies such as normalized RMSE (NRMSE), cosine similarity, and flow angular error. We evaluated post-training quantization of Cellpose–SAM, a promptless Segment–Anything model for cellular instance segmentation, on a stratified 245-field panel spanning BBBC038, BBBC039, and NIST iPSC images across density regimes. These proxies failed to reliably predict downstream mask retention under activation quantization. Angular error saturated at 90◦ from A8 through A2, while cosine and NRMSE preserved rank-order information without providing an interpretable acceptance threshold. At W8A8-cal-QDQ, cosine similarity was 0.944 despite 89/245 empty masks; moreover, performance varied sharply by modality, with BBBC039 yielding 0/47 empty masks versus 63/99 for NIST iPSC fields at the same proxy footprint. Across five modality–density strata, cosine similarity and empty-mask rate showed weak rank agreement (ρ≈0.3), with the highest-cosine stratum producing the second-worst downstream outcome. Below A8, all modalities collapsed to 245/245 empty fields. By contrast, ternary weight quantization produced a 12.08×storage reduction and was correctly identified as catastrophic by the proxies. These results show that FP32-referenced output proxies are insufficient as standalone acceptance criteria for promptless dense-prediction models; compression should instead be validated against downstream retention stratified by deployment-relevant imaging conditions.