NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Abstract
Robot fine-tuning can improve manipulation success while degrading the semantic structure inherited from large vision-language pretraining. This trade-off limits the reuse of vision-language-action (VLA) policies in settings that require compositional instructions, camera changes, and coordination with higher-level agents. We present \textbf{NoTVLA}, a semantics-preserving robot adaptation framework that replaces dense low-level action supervision with a sparse narrative action interface. NoTVLA converts demonstrations into sparse, semantically meaningful waypoints, grounds each decision with a task-relevant visual anchor and depth value, and reconstructs executable motion through a deterministic detokenizer. The resulting interface keeps the autoregressive vision-language backbone close to its pretrained prediction format while delegating high-frequency control to a transparent motion-rendering stage. We evaluate this design through matched and backbone-family comparisons, semantic retention probes, semantic out-of-distribution manipulation tasks, camera and depth perturbations, and deployment-oriented efficiency analysis. The results suggest that robot adaptation should be judged not only by task success, but also by how much task-relevant semantic competence is preserved during fine-tuning.