How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
Abstract
Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations: datasets at modern scale require thousands of skilled-operator hours on dedicated robot hardware. The language description paired with these demonstrations does not face the same bottleneck---a single short task label is cheap, but it also leaves implicit spatial relations, object interactions, embodiment state, and subgoal structure that the pixels already contain. We treat \emph{language density} as a cheap lever for amplifying signal in a fixed demonstration corpus; whereas prior dense-language work in robotics commits to a single caption style, we instead ask which kind of dense language helps each task and learn to deliver it at deployment. We realize this as \textbf{DeMiAn} (\textbf{De}nse \textbf{M}ult\textbf{i}-aspect \textbf{A}nnotatio\textbf{n}) in two stages. First, an automatic VLM pipeline re-labels each segment of an existing demonstration along four complementary aspects---\emph{physical motion}, \emph{scene composition}, \emph{arm pose}, and segment-level \emph{reasoning}---each surfacing a distinct kind of structure that a one-line task label omits. Second, a small learned \emph{instructor}, trained via supervised fine-tuning, maps the natural language task description and an initial scene snapshot to a task-appropriate annotation and runs asynchronously alongside the action policy, hiding generation latency behind the rollout. Applied to over 1M robot manipulation and 50K EgoVerse human-egocentric videos, with no new demonstrations collected, DeMiAn delivers four findings on a VLA action policy and a video-based world-action model: i) the learned instructor lifts RoboCasa success by 5 points over the no-annotation baseline, within 3 points of a per-task oracle; (ii) that oracle is non-trivial---no fixed aspect dominates, and peak performance requires selecting the right annotation aspect per task; (iii) the trained system extends usefully to composite tasks under subgoal-driven prompt switching, and to OOD scenes and objects; and (iv) dense annotation improves the compute-performance frontier in both mid-training and post-training, making re-annotation a practical scaling lever for robot policy learning.