StreamSE: Neural Audio Codec-based Streaming Speech Enhancement for On-device Inference
Abstract
Recent generative speech enhancement (SE) models achieve high perceptual quality, but their non-causal processing and expensive computation make them difficult to use for streaming, on-device inference. We propose StreamSE, a 20M-parameter streaming SE model that operates in the latent space of a frozen pretrained neural audio codec. At each frame, StreamSE directly regresses a clean continuous latent from the noisy codec latent, while conditioning on its own predictions from previous frames. This autoregression requires only a single forward pass per frame, with no iterative sampling or look-ahead. We further truncate the codec’s residual vector quantizer (RVQ) to its coarse levels, reducing representation complexity while improving robustness to reverberation. Although it is the only streaming generative model in our comparison, StreamSE reaches the perceptual quality of the offline generative baselines. Furthermore, the complete codec–StreamSE pipeline runs in real time on a single smartphone CPU core, with StreamSE accounting for only 10% of the per-frame latency. These results show that continuous latent regression provides a practical route to generative speech enhancement under fully causal, on-device streaming constraints.