StainNFT: Curriculum-Gated Multi-Reward Post-Training for Pathology-Faithful Virtual Staining
Abstract
Immunohistochemical (IHC) staining encodes molecular protein expression critical for clinical diagnosis, yet its chemical procedures are costly and time-consuming. Virtual staining offers a compelling alternative by digitally synthesizing IHC images from hematoxylin-eosin (H&E) stained slides. Despite remarkable progress in diffusion- and flow-matching-based virtual staining, supervised fine-tuning (SFT) remains fundamentally limited in pathological fidelity due to its coarse, spatially averaged supervision signal that starves sparse DAB-positive regions of effective gradients. Reinforcement learning (RL) can exploit inter-rollout variance for outcome-level pathological supervision, but naively applying DAB-based rewards triggers severe reward hacking that collapses the image distribution and degrades perceptual quality. This paper presents StainNFT, a flow matching RL post-training framework built on a curriculum reward strategy that gates fine-grained optical density supervision on per-sample DAB mask IoU. To enhance fine-grained pathological fidelity, we introduce expression-aware reweighting, multi-scale block supervision, and a closed-loop implicit pathological semantic reward derived from pathology foundation model priors. Extensive experiments on seven benchmarks demonstrate that StainNFT consistently outperforms existing methods in both perceptual quality and pathological fidelity for H&E-to-IHC virtual staining. Ablation studies confirm the effectiveness of each proposed component. Our code and trained models will be released.