Presence Is Not Placement: Auditing What Vibrato Metrics Measure
Abstract
Not all vibrato is created equal. Singers decide which notes receive more of it, and this decision contributes to phrasing. In VocalSet [1], notes held twice as long as their neighbors have about 16% deeper vibrato, and phrase-final notes have about 17% deeper vibrato. We ask whether a learned vibrato evaluator can detect this placement. To test this, we build WaveMaster, a MERT-based vibrato-presence classifier following the evaluation used in recent singing-synthesis work. Recordings are then edited note by note. Our main modification keeps the same set of vibrato depths but assigns those depths to different notes. This removes most of the measured placement while holding overall depth fixed. The classifier’s mean vibrato probability changes by only +0.001 (95% CI [−0.003, +0.005]). Giving every note the same vibrato depth also receives perfect Style Accuracy. The result replicates on eight singers recorded outside VocalSet and across four training variants. Intermediate MERT representations retain information about note depth, but the final time-averaged classifier does not use it effectively. This is not a failure of binary classification. It is a mismatch between what the metric measures and what expressive-control systems aim to preserve. A synthesizer optimized only against this presence metric could therefore discard phrasing without penalty. Vibrato controls should be evaluated not only by whether it is present, but also by where it occurs.