GIVLA: Deep Geometry Internalization for A Lightweight VLA via Geometry Instruction and Gradient-Informed Training
YUNHE LI ⋅ Qiming Liu ⋅ Haoyuan Wang ⋅ Hesheng Wang
Abstract
Vision-Language-Action (VLA) models often struggle with precise manipulation due to their lack of explicit spatial awareness. To bridge this gap, we propose GIVLA, a framework that internalizes geometric priors into the VLA backbone through a deep-coupled architecture. GIVLA implements a two-stage gradient-informed training paradigm to resolve task-level interference and ensure precise execution, a design grounded in the perception-action duality identified through our analysis of training dynamics. Extensive experiments on LIBERO and physical robot platforms demonstrate that GIVLA achieves superior accuracy and robustness with high parameter efficiency.
Chat is not available.
Successful Page Load