Efficient Brain-to-Speech Decoding with Fixed-Delay Spiking Neural Networks
Aleksandra Wisniewska ⋅ Seo-Hyun Lee ⋅ Seong-Whan Lee
Abstract
Brain-to-speech decoding aims to restore communication by translating neural activity into speech-related representations. From a modeling perspective, this task can be formulated as a weakly supervised neural sequence decoding problem, where phoneme sequences must be recovered without frame-level annotations. Existing speech decoders commonly rely on recurrent sequence models to integrate temporal context, but their dense sequential computation can limit their suitability for resource-constrained portable or implantable neuroprosthetic systems. In this study, we propose a residual spiking neural network (SNN) for phoneme-level decoding of intracortical speech signals under constrained model capacity and weak sequence-level supervision. Our approach incorporates biologically inspired fixed connection-specific delays to integrate temporal context, providing each postsynaptic unit with access to recent presynaptic spike activity without recurrent state or dense temporal convolutions. Because each directed connection is assigned a single fixed delay, the temporal field is expanded without adding trainable delay parameters. Across 500k, 2M, and 5M core-parameter budgets, fixed delays consistently improve phoneme error rate (PER), with gains saturating once the delay range covers the task-relevant temporal horizon. Boundary-aligned CTC analysis further suggests that delayed temporal context reduces local uncertainty around phoneme transitions, indicating that fixed delays support weakly supervised phoneme alignment rather than simply increasing model capacity. Compared with recurrent, convolutional, and feedforward baselines under matched core-parameter budgets, the proposed SNN achieves the best PER at 500k and 2M parameters and the second-best PER at 5M parameters, while reducing estimated operation-level energy by approximately 2-4$\times$ relative to the strongest non-spiking baselines. These results suggest that fixed-delay SNNs provide an efficient temporal modeling mechanism for brain-to-speech decoding, potentially enabling lightweight decoding for on-device speech neuroprostheses.
Chat is not available.
Successful Page Load