Pixels to Tokens: Token Space Efficient Active Learning for Low-Budget Semantic Segmentation
Utku Cicek ⋅ Mahip Singh ⋅ Ziyao Shang ⋅ Kimathi Kaai ⋅ C Thomas ⋅ Pablo Guerrero ⋅ Alexander Wong ⋅ Sirisha Rambhatla
Abstract
Semantic segmentation typically requires dense pixel-level annotations, creating a prohibitive bottleneck for budget-constrained applications. While foundation models and vision transformers (ViTs) have redefined visual representations, current active learning (AL) strategies for these backbones largely operate at the patch or region level, where annotation requirements remain high. Extending ViT-based AL to the pixel-level, low-budget regime is uniquely challenging: with as little as one pixel query per image, a method must simultaneously identify the most informative locations and train a reliable decoder head from an extremely sparse signal. We introduce TEAL (Token-space Efficient Active Learning), the first pixel-level AL framework built around ViT representations. TEAL uses a frozen DINOv3 backbone and performs diversity-based selection directly on its native token lattice, avoiding the artifacts introduced by interpolating token embeddings to dense pixel resolutions. Our framework applies a two-level MaxHerding strategy over multi-layer token descriptors to select representative candidates, followed by margin-based uncertainty refinement on a finer decoder grid. Across CamVid, Cityscapes, ADE20K, and Pascal Context, TEAL consistently outperforms previous baselines under extreme label scarcity. After 10 rounds of 1-pixel-per-image queries, TEAL improves over the strongest baseline by up to $+\textbf{14.65}$ mIoU on CamVid and $+\textbf{23.34}$ mIoU on Pascal Context, showing that ViT token spaces provide an effective geometry for extremely-low-budget active segmentation.
Chat is not available.
Successful Page Load