Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Abstract
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for embodied AI. While previous research primarily leverages static cues from language or single images, which often lack the rich interaction context necessary to accurately locate functional zones. To alleviate this predicament, we collect a comprehensive video-based 3D affordance dataset, VIDA, which contains 38K human-object-interaction (HOI) videos covering 16 affordance types, 38 object categories, and 22K point clouds. Based on VIDA, we propose a strong baseline: VideoAfford, a unified framework that extends Multimodal Large Language Models (MLLMs) with fine-grained affordance segmentation capabilities. VideoAfford incorporates a latent action encoder to distill dynamic interaction knowledge from demonstration videos and a spatial-aware loss to encourage geometrically consistent affordance predictions. Extensive experiments on VIDA show that VideoAfford significantly outperforms strong baselines in both seen and unseen settings, demonstrating its effectiveness in video-driven 3D affordance reasoning and open-world generalization. The dataset and code will be released soon to facilitate future research in this area.