BoxTuning: Object-Aware Visual Prompting for Multimodal Model Fine-Tuning
Abstract
Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack explicit mechanisms for fine-grained object grounding. Recent work addresses this by serializing bounding box coordinates as text tokens, but this text-coordinate paradigm suffers from a fundamental modality mismatch: object information is inherently visual, yet encoding it as text incurs a high token cost that forces aggressive temporal downsampling. We propose BoxTuning, an object-aware visual prompting framework for multimodal model fine-tuning that moves spatial-temporal object states from text-coordinate serialization into the visual stream while retaining only object identity in a compact text legend. Colored bounding boxes and trajectory trails encode geometry and motion on video frames, while the color-to-object legend provides the minimal textual identity bridge. This reduces the token cost significantly, achieving 87-93\% text token reduction in practice. It also preserves full temporal resolution, where the trajectory trails further encode inter-frame motion direction and speed within each keyframe, recovering fine-grained dynamics that text-coordinate methods are forced to discard. Experimental results on five video QA benchmarks (CLEVRER, Perception Test, STAR, NExT-QA, IntentQA) show that BoxTuning surpasses text-coordinate baselines on spatially oriented tasks and avoids their degradation on reasoning-centric tasks, establishing object-aware visual prompting as a natural and efficient way to convey object information to MLLMs.