Mind the Gap: Dataset and Fine-grained Evaluation for Inline Audio Descriptions
Subhashini Venugopalan ⋅ Yingwen Tan ⋅ Taylor Roper ⋅ Jimmy Tobin ⋅ Anton Kast ⋅ Alicia Martin ⋅ Sam Sepah ⋅ Amy Pavel
Abstract
Audio descriptions (AD) are essential for making visual media accessible to blind and low-vision (BLV) users. While Multimodal Large Language Models (MLLMs) offer a scalable solution for generating audio descriptions on-demand, their performance on ``inline'' audio descriptions -- which must fit within existing silences in a video -- remains under-explored, particularly for diverse user-generated content (UGC). We present a comprehensive investigation into MLLM-generated inline audio descriptions. First, we introduce a professionally annotated dataset of 36 hours of videos, each averaging $\sim$5 mins. and total 7k+ descriptions, providing a high-quality benchmark for this task. Second, we propose an evaluation framework based on audio description expert guidelines that provides fine-grained actionable metrics for model improvement. Crucially, our framework evaluates the end-to-end task: it renders the generated scripts into audio tracks to assess their fit within the original video's natural silences, minimizing disruptive overlaps. Our analysis reveals that while frontier MLLMs accurately describe visual events, significant gaps persist in audio-visual timing, quality, and narrative flow. Finally, we validate the framework and conduct studies with BLV participants and sighted raters, finding that while users narrowly prefer human-authored audio descriptions, they still find MLLM-generated audio descriptions highly beneficial for comprehension. We provide guidance on the current capabilities of MLLMs for audio descriptions and release our dataset and evaluation suite to the community.
Chat is not available.
Successful Page Load