Breaking the Second Barrier: Sub-Second Timestamped Omni-Modal Captioning
Abstract
Omni-modal caption models equipped with precise timestamp grounding are crucial for fine-grained video understanding and temporally controllable video generation. However, existing open-source models, and even proprietary systems, struggle to provide fine-grained timestamp grounding for short video captioning. To bridge this gap, we present Deci-Omni-Captioner, an advanced open-source omni-modal captioning framework designed for sub-second timestamp precision. During the Supervised Fine-Tuning (SFT) stage, we leverage Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models to strictly calibrate visual text and spoken dialogue boundaries, offering highly deterministic temporal anchors to endow the model with sub-second timestamp grounding capabilities. Furthermore, we curate a specialized dataset for multi-reward Reinforcement Learning (RL) and propose Span-Specific Credit Assignment (SSCA). Unlike conventional global advantage normalization, which dilutes reward signals across weakly coupled multimodal descriptions, SSCA calculates advantages independently for different structural spans, effectively isolating penalties and rewards across semantic and temporal dimensions. Extensive experiments demonstrate that Deci-Omni-Captioner establishes state-of-the-art (SOTA) performance among open-source models on multiple semantic comprehension benchmarks (e.g., AVUT, UGC, VCapsBench). On our challenging Deci-Timestamp Bench, it significantly surpasses leading proprietary models, including the Gemini 2.5 and 3.0 series.