Unlocking Region-Level Understanding by Steering Frozen VLMs
Abstract
Large Vision-Language Models (LVLMs) are primarily trained for generative objectives and inherently lack a mechanism to produce precise, localized embeddings for discriminative tasks. Bridging this gap typically requires fine-tuning the massive backbone, which is computationally expensive and risks catastrophic forgetting. We propose a parameter-efficient method to {unlock} region-level understanding in frozen VLMs, allowing a single model to serve as both a capable generator and a high-precision region embedder. By training only appended soft tokens and a learnable layer aggregation mechanism (updating only 0.001\% of the vision backbone's parameters) alongside a fully-trained, decoupled 270M text encoder, we steer the frozen backbone to produce high-quality region embeddings. We introduce a prompting strategy that directly injects visual region features (RoI tokens), which significantly outperforms standard text-coordinate prompting. Using the publicly available Gemma-4B models, we show that our highly efficient method significantly outperforms LoRA-adapted models and even massive generative baseline Gemma-27B on region-level classification tasks. Furthermore, our final model surpasses the performance of prior fully-finetuned region-level CLIP methods, highlighting an effective application of our parameter-efficient method for open-vocabulary region-level representation and classification.