Which singing technique degrades first? Technique contrast retention for neural audio codecs
Abstract
Neural audio codecs tokenize singing voice, but overall quality scores do not show which vocal techniques survive. We introduce Technique Contrast Retention (TCR), measuring how much of the contrast between matched technique and plain recordings of the same singer and phrase survives compression. TCR is a dimensionless, group-level retention slope, accompanied by absolute contrast error and direction consistency, and reported alongside reconstruction quality. On Chinese and English GTSinger recordings, per-technique linear probes on a frozen self-supervised representation measure six techniques. Across six codecs, we map technique fragility; within codecs, bypass and rate interventions distinguish quantization-associated losses from continuous-latent-path losses. A fully nested re-estimation preserves all 72 monotone rate steps, while exact cross-technique ordering is selection-sensitive. A listening study on vibrato, breathy, and glissando (24 of 25 screened listeners retained) finds audible differences graded in the retention gap (pooled model-based agreement 0.95; raw agreement 623/696 ≈ 0.90). The other three techniques are not tested perceptually.