Causal-VLM: Dense Causal Captioning in Videos
Abstract
Dense video captioning describes multiple events in long videos. However, each event is predicted independently, ignoring the causal relationships with the other events in the video. These causal relationships are important for human-like descriptions of long videos. Our experiments show that exisiting vision-language models fail at predicting causal relationships in long-form videos. No prior work jointly generates captions and event timestamps and predicts causal relationships over multi-event sequences in long-form videos. We address this through two contributions. First, we create a benchmark dataset with 21.6K videos and 85K events by using a large language model to generate annotations from dense captions in YouCook2 and ActivityNet. We validate this benchmark dataset through multimodal alignment, vision-language models, and human judgment. Second, we propose a novel architecture for dense causal captioning that not only generates captions but also learns causal relationships. This novel architecture not only predicts causality effectively (0.73 F1 ActivityNet, 0.64 F1 YouCook2) but also improves the vision-language model's core capabilities: caption quality increases by +6.3 CIDEr on ActivityNet and +7.4 on YouCook2, while temporal grounding becomes accurate as causal predictions force precise event boundary localization.