FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Abstract
Vision-Language-Action (VLA) models have advanced rapidly by leveraging large-scale vision-language priors and robot trajectory data, yet they remain weak at following fine-grained human instructions. We argue that this limitation stems from a fundamental mismatch between language and action supervision in existing robot datasets. Most open-source datasets annotate each trajectory with a goal-oriented task description, while leaving the underlying execution process largely unspecified, including motion trajectories, contact and approach patterns, object state transitions, recovery behavior, and final configurations. Since different choices along these dimensions can lead to substantially different action sequences, coarse task descriptions induce a many-to-one mapping from action trajectories to language, preventing precise action-instruction alignment. To address this problem, we propose FineVLA, a framework for constructing and leveraging fine-grained embodied action instructions for VLA learning. FineVLA includes FineVLA-Tool, a scalable pipeline for data cleaning, temporal alignment, and fine-grained annotation; RoboFine-VLM, a vision-language model for fine-grained robotic action understanding and annotation; RoboFine-Bench, a benchmark for evaluating fine-grained robotic action understanding via VQA and captioning; and FineVLA-Policy, a VLA policy trained with fine-grained instructions on open-source ALOHA robot data. Across benchmark evaluation and policy learning experiments, FineVLA demonstrates that high-density, action-aligned language supervision leads to more controllable and instruction-sensitive robot policies. Our results highlight fine-grained instruction alignment as a critical step toward VLA systems that not only complete tasks, but complete them in the way humans specify.