Compute-Optimal Scaling Laws for Spiking Language Models
Abstract
Spiking neural networks target energy-efficient language modeling through sparse, binary activations, but the efficiency–accuracy tradeoff of spiking language models is uncharacterized at scale. We present the first Chinchilla-style IsoFLOP scaling study of a spiking language model, sweeping SpikeGPT-family models (RWKV backbone) from width 384 to 1536 across compute budgets from 3×10^17 to 1×10^19 FLOPs on FineWeb-EDU. We report two findings. First, the compute-optimal frontier is N_opt ∝ C^0.56, close to the 0.50 that Hoffmann et al. report for dense Transformers, so spiking does not change how parameters and data should be traded off under a compute budget. Second, at matched compute, binary activations cost roughly 0.3 nats of validation loss relative to a non-spiking RWKV baseline. We further track how activation sparsity evolves with model size and training. These results give the first compute-optimal recipe for spiking language models and quantify the accuracy cost of their sparsity.