MoRe-DVC: Motion Retrieval-Augmented Generation for Detailed Video Captioning
Abstract
Despite significant advances in large vision-language models (Video-LLMs) for general video understanding, accurately narrating fine-grained, highly dynamic human activities remains a formidable challenge. Existing approaches typically rely on global sequence embeddings or isolated part-level retrieval, lacking the precise relational kinematic grounding required to prevent physical hallucinations. To address this, we introduce Motion Retrieval-Augmented Generation for Detailed Video Captioning (MoRe-DVC). Our algorithmic novelty centers on two core contributions: an event-boundary-aware graph-conditioned retrieval mechanism that captures hierarchical and chronological motion dependencies, and a symbolic physical faithfulness reward that explicitly constrains generation to grounded physical reality. Specifically, MoRe-DVC integrates a Pose-Orientation Encoding Module (POEM) to extract symbolic descriptors, an Event Calibration and Refinement Module (ECRM) to dynamically adjust coarse action boundaries into an event graph, and a Graph-Conditioned Hybrid Retriever (GCHR) to select relationally relevant evidence. By coupling graph-level retrieval with a strict symbolic reward, our framework reduces motion-language inconsistency. Extensive experiments demonstrate that MoRe-DVC achieves robust outperformance across three motion-centric datasets (BoFiT, HumanML3D, and FineMotion), establishing a new standard for physically faithful action narration.