Depth through Recurrence: Towards Ultra-Efficient On-Device ASR
Chen Feng ⋅ Tianyi Xu ⋅ Yicheng Lin ⋅ Jay Zhuo ⋅ Ramchalam Kinattinkara Ramakrishnan ⋅ Zhaocong Yuan ⋅ Chenzheng Su ⋅ Xiaopeng Zhang
Abstract
Modern automatic speech recognition (ASR) systems have achieved remarkable accuracy by scaling model depth and capacity, but at the cost of substantial memory and computation. On edge devices where ASR is often most needed, such as watches and glasses, deploying such large models is infeasible due to extremely constrained resources. This raises a fundamental question: can we achieve high representational power in deep ASR models without scaling up parameterization? In this work, we revisit the role of depth and identify layer-wise representational dynamics, in which most layers learn functionally similar transformations. Motivated by this insight, we propose a block-recurrent ASR architecture that replaces parameterized depth with a small set of recurrent blocks, each consisting of weight-shared layers, thereby preserving effective depth while drastically reducing model size. Through representation-guided grouping and knowledge distillation, block-recurrent models retain the accuracy of large reference models while using an order of magnitude fewer parameters. Across Open ASR Leaderboard benchmarks, a model with only two shared blocks recovers $97.4$% of the accuracy of large models on average. The compact backbone further enables efficient on-device personalization through lightweight MoE-LoRA adaptation, allowing user-specific models to match or surpass the accuracy of general-purpose ASR systems with significantly lower inference cost. Together, these results show that high-quality ASR can be achieved without large parameterization, and that structured recurrence provides a principled path towards ultra-compressed, high-efficiency, and personalized speech recognition on the edge.
Chat is not available.
Successful Page Load