Factorized Asymmetric Conditioning for Efficient Transport in High-Fidelity Fluorescence Microscopy Synthesis
Abstract
The spatial phenotype of a protein emerges from the interaction between its molecular identity and the surrounding cellular environment. Synthesizing fluorescence microscopy images from protein sequence and cellular morphology is therefore an asymmetric multimodal generation problem: morphology provides spatially aligned structural information, whereas sequence provides protein-specific semantics. We introduce FACET (Factorized Asymmetric Conditioning for Efficient Transport), a probabilistic framework that assigns these sources of variation distinct but coupled roles. FACET separates context-explainable structure, coarse protein semantics, fine protein-specific variation, and image stochasticity through context-residualized semantic memory, a bounded variational residual, and context-first hierarchical conditioning. It couples this representation with stochastic flow transport and a variance-preserving projection that enables efficient adaptation of a pretrained discrete-time diffusion predictor. Experiments on HPA and OpenCell demonstrate superior spatial-phenotype agreement and distributional fidelity at substantially lower sampling cost. Controlled analyses further show robustness across cellular context compositions, stronger alignment between generated phenotypes and protein-interaction structure, and complementary contributions from the semantic, hierarchical, and transport components. Together, FACET provides a general framework for high-fidelity fluorescence microscopy synthesis when multimodal conditions differ in structure, alignment, and predictive strength.