Windchime: an Audiovisual Installation with Audio-Language Model implementations
Abstract
Windchime is an audiovisual installation in which spoken language retrieves music and interactive animations from an artist authored corpus. A local Automatic Speech Recognition (ASR) model transcribes each prompt, an audio-language model (ALM) embeds it and retrieval over 406 audio stems returns a mix with live coded Strudel fragments that layer into audio output. Simultaneously each visitor’s input results in retrieval of interactive Three.js animation scenes each containing unique tactile control via monome grid and arc hardware. Windchime distributes agency in three distinct directions; across the artist who authored the possible outputs, the ALM that influences what areas of the corpus can be retrieved, and the visitor who’s spoken inputs result in multimodal embodied interaction.