Parameter-Efficient Objective Design for Masked Autoencoding of Multiplexed Tissue Images
Adriano Martinelli ⋅ Marianna Rapsomaniki
Abstract
Masked autoencoders (MAEs) have been used to learn general-purpose representations from natural images and are increasingly being applied to biomedical modalities such as imaging mass cytometry (IMC). MAE training, however, was developed for natural images, whereas multiplexed tissue images have markedly different statistical properties. Each IMC channel measures a distinct molecular marker, and marker signals vary widely in prevalence and dynamic range. We test two adaptations to the MAE objective: an auxiliary cross-marker correlation loss ($L_\text{corr}$) that predicts tile-level cross-marker correlations and a weighting scheme that directs more of the reconstruction gradient towards signal-rich patches. We evaluate representations across three cancer types using tile-level cell type abundance regression and, in two cohorts, patient-level classification of clinical endpoints comprising subtype, grade and ER status. Under common downstream readouts, we compare against image-statistic baselines, untrained encoders, and VirTues, a recently published foundation model. Adding the correlation objective yields higher five-fold average scores on every primary endpoint at every matched capacity. A simple channel-correlation baseline, however, achieves higher abundance $R^2$ than vanilla MAE in two of the three cohorts and than VirTues on the tested cohort, showing that encoder pretraining does not by itself outperform simple image statistics. Model-scaling experiments show that greater encoder capacity improves abundance $R^2$ while decoder capacity has smaller, more variable effects; more broadly, capacity scaling improves performance less consistently and to a smaller extent than the correlation objective. Importantly, even the smallest $L_\text{corr}$-trained model configuration achieves higher average scores than every reconstruction-only configuration across all six primary endpoints; the best $L_\text{corr}$-trained configuration also achieves higher averages than VirTues on both evaluated endpoints. Overall, adapting the training objective improves model performance more efficiently than increasing model capacity and demonstrates the importance of image-statistic and untrained encoder controls when evaluating spatial-omics representations.
Chat is not available.
Successful Page Load