Dual-Contrastive Sparse Autoencoders Reveal Features of Musical Interpretation
Abstract
For decades, philosophers and musicologists have debated which features of a performance are constitutive of the work and which are expressions of how it is being played. Computational evidence has been hard to assemble: audio embedding models conflate work identity with performance style, and existing interpretability tools for music generators treat learned representations as flat dictionaries that mix the two. We turn the question into an empirical one by introducing the dual-contrastive sparse autoencoder (DC-SAE), a two-branch sparse autoencoder that uses cheap work-level metadata to factor a generative music transformer's residual stream into work-identity and performance-content subspaces. Across jazz-standard, classical-work, and pop-cover corpora, the resulting decomposition is supported by probes and feature galleries that surface musically interpretable concepts on each side. Without performer supervision, the variation branch acquires structure related to performer identity, and steering along these directions can shift the perceived performer of generated audio while preserving the underlying work. Together, these results show that a generative music model has internalized a representation of interpretation that can be recovered, named, and steered.