RRL-HOI: Reflective Reinforcement Learning for Open-Vocabulary HOI Detection
Abstract
Open-Vocabulary Human-Object Interaction (HOI) detection aims to localize interacting human-object pairs while generalizing to novel interaction categories beyond the training set. Although Multimodal Large Language Models (MLLMs) exhibit strong open-vocabulary understanding, directly applying MLLMs to HOI detection remains challenging, as language-prior hallucinations and coupled localization-semantic errors can destabilize structured HOI predictions. To address these issues, we propose RRL-HOI, a novel Reflective Reinforcement Learning framework that trains MLLMs to follow a verify-and-revise self-correction policy. Specifically, RRL-HOI first adapts MLLMs to produce box-grounded HOI triplets, and then introduces an executable reflection mechanism that targets the two failure modes through structured edits over localization, human-object pairing, and interaction semantics. In this way, localization and pairing edits mitigate localization-semantic coupling, while interaction-semantic edits suppress visually unsupported language-prior predictions. To further improve this reflection mechanism, RRL-HOI formulates reflection as a reinforcement learning objective trained with GRPO, where the reward jointly considers localization quality, pair-association correctness, interaction correctness, and edit consistency. Through reward-driven policy learning, RRL-HOI turns open-vocabulary recognition from a single-step generative prediction into an explicit verify-and-revise decision process, thereby bridging the gap between general multimodal understanding and precise interaction detection. Extensive experiments on two standard open-vocabulary HOI detection benchmarks demonstrate the effectiveness of our method.