Do Speech BCIs Need Larger Models? Rethinking Neural Decoding beyond Scaling
Zehui Feng ⋅ Weichuan Wang ⋅ Xiaohan Chen ⋅ Cuntai Guan ⋅ Ting Han
Abstract
Recent speech brain-computer interfaces (BCIs) increasingly rely on large models and complex multi-stage pipelines for neural decoding. However, in speech BCI, where training data are scarce, noisy, and expensive to collect, the benefits of model scaling remain unclear. This work revisits a fundamental question: $\textbf{are large models truly necessary for accurate neural decoding?}$ We propose $\textbf{\textit{BEST}}$ ($\textbf{\textit{B}}$rain-$\textbf{\textit{E}}$ncoding-$\textbf{\textit{S}}$peech-$\textbf{\textit{T}}$ext), a two-stage framework for efficient neural-to-text decoding. Instead of relying on large generative models or phoneme-centric pipelines, $\textbf{\textit{BEST}}$ introduces hierarchical intermediate representations and a compact speech-oriented decoder that directly maps neural activity to text. This design preserves essential structure while significantly reducing computational cost, enabling robust generalization in low-resource settings and practical deployment under constrained resources. On the Brain-to-Text ’24 and ’25 benchmarks, $\textbf{\textit{BEST}}$ achieves word error rates (WERs) of 6.19\% and 1.62\%, espectively, using only 3.7\% of the parameters required by comparable methods. Beyond these results, we provide three key insights: (1) decoder scaling yields non-monotonic gains; (2) cross-modal alignment is more critical than model size; and (3) overfitting degards sequence-level performance. Representation analysis further reveals a consistent neural→audio→text alignment pathway.
Chat is not available.
Successful Page Load