Exploring Textual Guidance for Video-to-Music Generation Models
Vaibhavi Lokegaonkar ⋅ Aryan Vijay Bhosale ⋅ Vishnu Raj ⋅ Gouthaman KV ⋅ Ramani Duraiswami ⋅ Lie Lu ⋅ Sreyan Ghosh ⋅ Dinesh Manocha
Abstract
Current video-conditional music generators expose a narrow control interface, leaving the creator unable to specify the music a scored video should carry. The same clip is consistent with many combinations of genre, instrumentation, key, and tempo, and only the creator's intent selects among them. This also goes unmeasured, since the field reports audio quality and audio-visual similarity but not whether a stated attribute was realised. We explore this gap with \textbf{ReelBench}, a benchmark for the text$+$video-to-music task of 300 video, music, and composite text instruction triads annotated for tempo, key, instrumentation, and genre. We find that instruction adherence and audio quality are not achieved together by any system we evaluate, and that adherence is consistently lower for key than for tempo. We then ask whether the diffusion autoregressive (DAR) paradigm, so far applied only to speech and song synthesis, transfers to this setting, and develop \textbf{Video-Robin} by attaching a video adapter to a patch-based DAR music generator. It obtains the lowest FAD in our comparison at the shortest inference time, on 1.0 B parameters and 311 hours of paired data, while following tempo close to the strongest baseline. We highlight the need to measure and improve instruction adherence in video-to-music generation.
Chat is not available.
Successful Page Load