Drive-to-Music: Grounded Context-Aware Music Generation with Multimodal Safety for Vehicles
Abstract
Music shapes driver attention, mood, and behavior, yet in-vehicle audio remains largely static and context-blind. Generative AI could enable more adaptive experiences, but its deployment in vehicles raises safety challenges that content moderation alone cannot address. We present Drive-to-Music, a context-aware system that generates music and complementary visual elements grounded in multimodal driving signals. Using dash-cam imagery, vehicle telemetry, and environmental information, the system derives a grounded representation of the driving context, translates it into musical descriptors, and generates a contextually aligned soundtrack, cover art, and track title. To support safe in-vehicle deployment, we introduce a multimodal contextual safety framework that evaluates the generated music, cover art, and title both independently and in relation to the current driving context. The framework combines modality-specific risk estimates with cross-modal consistency analysis to identify direct harms and contextual risks, such as urgency-inducing content during high-speed or low-visibility driving. Constraint-based controls then accept, adapt, or regenerate content before it reaches the driver. Drive-to-Music provides a foundation for grounded, responsible, context-adaptive generative AI experiences in vehicles.