Sequential Set Prediction from Stochastic Fluorophore Blinking: Supervision and Architecture in SMLM
Abstract
We study post-processing of single-molecule localization microscopy (SMLM) acquisitions generated by stochastic fluorophore blinking. The final reconstruction uses the complete recorded acquisition, while our causal models process localization events sequentially and can update the predicted molecule set after each event. We ask how these sequential observations should be supervised during training. Across Mamba-2 and Transformer backbones, model widths, and short- and long-dark-time conditions, dense causal/evolving supervision improves localization in every configuration and molecule detection in nearly all configurations relative to a single terminal target. The effect is substantially larger for the Transformer and, for that backbone, is comparable in magnitude to changing the sequence architecture. With supervision fixed, the Transformer achieves better molecule detection than Mamba-2, while localization accuracy remains similar. The long-dark-time condition shows lower molecule detection in nearly all settings, while localization is preserved or improved in most cases. Thus, in this stochastic observation setting, the supervision interface is a first-order design choice whose impact depends on the sequence backbone.