Speech Emotion Recognition with Classical and Deep Learning Models: An Accuracy–Efficiency Frontier from MFCC to Self-Supervised Representations
Abstract
Speech Emotion Recognition (SER) increasingly relies on large self-supervised representations, but their accuracy gains over lightweight hand-crafted pipelines are rarely quantified against real deployment cost. We map this trade-off directly: an MFCC-based pipeline (SVM, MLP) with Gradual Magnitude Pruning, INT8 quantization, and K-Means diversity coresets is compared against frozen wav2vec2-base embeddings on RAVDESS and CREMA-D, measuring real (not theoretical) model size, CPU latency, and throughput alongside accuracy. To obtain accuracy numbers that reflect genuine generalization rather than evaluation artifacts, we replace the standard random stratified split, which allows speakers to appear in both training and test sets, with 5-fold speaker-independent (GroupKFold) cross-validation throughout. Frozen wav2vec2 embeddings outperform our tuned MFCC+MLP pipeline by 11.3 pp (RAVDESS) and 13.2 pp (CREMA-D), at a cost of 4,833x and 966x more parameters (377.6 MB, 94.4M vs. 0.08-0.40 MB, 19.5K-97.7K) - defining a clear accuracy-efficiency frontier between the two approaches. Measured INT8 quantization gives a real 2.82-3.58x size reduction (vs. a theoretical 4x) but no latency benefit over fp32, showing that compression ratio and inference speedup are distinct quantities. A notable methodological finding emerged from adopting the speaker-independent protocol: RAVDESS accuracy fell sharply under speaker-independent evaluation, a drop that is statistically significant for every model/dataset pair tested (p < 0.005, one-sample t-test vs. leaky baseline), while CREMA-D, with nearly 4x more speakers, was far more stable - indicating that standard stratified splits can substantially inflate reported SER accuracy, while suggesting that dataset speaker composition may affect benchmark reliability. Confusion matrices are consistent across datasets: high-arousal Angry is best recognized (0.66 recall, both datasets), while low-arousal emotions (Neutral, Sad, Calm) form a persistent confusion cluster. Together, these results position lightweight MFCC pipelines and large self-supervised models as two ends of an explicit accuracy-efficiency frontier for SER, grounded in an evaluation protocol robust to speaker-identity leakage.