DiffCap-RL: Differential QA Rewards for Dense Image and Video Captioning
Abstract
Dense image and video captioning demands exhaustive visual details while strictly maintaining factuality, forcing a severe trade-off between descriptive density and hallucination. Supervised fine-tuning (SFT) struggles here: for dense captions, its token-level loss dilutes effective signals while failing to reinforce specific fine-grained dimensions. Reinforcement learning (RL) is a promising alternative but remains bottlenecked by reward design: coarse-grained VLM judges suffer from reward drift, and static QA rewards quickly saturate as the policy improves. To address this, we propose DiffCap-RL, an RL framework combining a stable offline Pre-Scorer with an adaptive online Diff Scorer. We utilize the Pre-Scorer as a stable anchor for broad visual grounding via automated QA pairs. Once it loses discriminative power, Online Diff Scorer dynamically extracts semantic disagreements between policy rollouts. By decomposing these disagreements into targeted conflict questions to suppress hallucinations and extra questions to reward valid new details, the Diff Scorer provides reliable, non-saturating supervision at the policy frontier. We validate our reward quality from three perspectives: human alignment, online RL signal quality, and BoN selection. Across multiple image and video captioning benchmarks, our DiffCap-RL delivers large gains and consistently outperforms state-of-the-art baselines at similar or larger scales.