Adapting Vision Transformers to Organoid Imaging
Abstract
Organoids are an increasingly central scientific platform for drug screening, disease modeling, and regenerative medicine. Because organoid form is closely tied to function, AI-driven image analysis has drawn intense interest, yet every existing organoid imaging study evaluates representations on a single dataset and no work has quantitatively compared representations across the breadth of public organoid datasets. We close this gap with \textbf{OrgBench}, a benchmark of 12 vision foundation models (VFMs) on 13 public organoid-imaging datasets spanning detection, segmentation, classification, and temporal outcome prediction. The benchmark reveals a structured domain gap: task-family winners shift across the benchmark and parameter count is a poor predictor of transfer. Building on this, we introduce \textbf{OrgFM}, a small bottleneck adapter that we continually pre-train on unlabeled organoid images while keeping a general-purpose vision transformer frozen. On the OrgBench task-family aggregates, OrgFM improves over its frozen base backbone. Layer- and feature-level interventions localize the gains to a small set of late-layer morphology-sensitive features: the adapter sparsely supplements, rather than replaces, the base representation. Together, this is the first cross-task, cross-backbone study of organoid imaging at scale; the OrgBench/OrgFM pairing establishes that the parameter-efficient adapter direction is a promising path for organoid representations and clarifies why it works.