Controlled Generation of Multiplexed Images with Flow Matching
Abstract
Spatial proteomics enables the quantification of protein expression by aggregating pixel intensities at the cell level, an aggregation corrupted by \emph{spillover}: signal emitted by one cell is assigned to its neighbors due to segmentation errors or instrument limits. Progress is limited by the absence of datasets in which the true origin of the signal is known, so decontamination methods can only be trained and scored against annotations that are themselves contaminated. We propose to build a spatial proteomics generation method with a conditional flow matching model that maps Gaussian noise to real protein-image while being conditioned, at full pixel resolution, on the localization and the annotated positivity of every cell in the patch, classifier-free guidance controlling how strictly the generated signal adheres to that prescription. We also propose to evaluate this generative model not on pure similarity but on utility by training a positivity classifier on synthetic data and test it on a manually curated real dataset. We also provide external validation by testing the model capability to follow the conditioning, measure the realism of the generated patches and their diversity. We find that a classifier trained on synthetic slides alone reaches an AUC of 0.915 and F1-score of 0.571 compared to the same classifier trained on real patches reaching 0.917 AUC and 0.601 F1-score.