Diving-R1: Empowering Multimodal LLMs with Traceable Progressive Reasoning for Interpretable Diving Action Quality Assessment
Abstract
Action Quality Assessment (AQA) for diving aims to quantify the execution quality of diving sequences and has long been a prominent research topic. However, existing diving AQA methods typically regress solely the final diving score, leading to poor interpretability. To address this challenge, we propose Diving-R1, the first attempt to leverage a Multimodal Large Language Model (MLLM) for interpretable diving AQA, enabling the integration of comprehensive information, thereby supporting progressive reasoning and traceable assessment. We construct a novel dataset, DivingThink, as the foundation, which contains progressively structured reasoning chains that follow judging logic and explicitly describe the execution quality of each diving sub-action, providing clear evidence for score determination. Furthermore, we establish a new benchmark, DivingInterp, which evaluates the performance of MLLMs on interpretable diving AQA from multiple complementary perspectives, including score estimation, textual similarity, and reasoning quality. Additionally, we carefully design a three-stage training paradigm, combined with a dedicated composite reward function, which gradually relaxes the annotation requirements for the training corpus at each stage to incrementally enrich the diversity of the training data, achieving continuous improvement. Extensive experiments demonstrate the superiority of Diving-R1 over existing MLLMs in interpretable diving AQA. The dataset and code will be available on GitHub.