Comparing Transformers and Hybrid Models at the Token Level
Yanhong Li ⋅ Will Merrill
Abstract
Hybrid language models that mix attention and recurrent layers have shown promise: theoretically, recurrent layers ameliorate the limitations of pure transformers on state tracking, and empirically, hybrids can outperform pure transformers in loss and downstream evaluations \citep{waleffe2024empirical,merrill2026olmohybrid}. Yet it remains unclear which data or capabilities drive these gains, and to what degree they reflect the theoretical advantages motivating hybrid models. We address this by analyzing the benefits of hybrid models at the token level---which kinds of tokens are better predicted by hybrids than by pure transformers---using the open weights from Olmo 3 \citep{olmo2025olmo3} and Olmo Hybrid \citep{merrill2026olmohybrid}. Hybrids predict tokens better across the board, but we show these gains localize to open-class content words rather than closed-class function words. Across natural language and code, the loss gap shrinks on closing brackets (but not opening brackets), consistent with the hypothesis that attention, not hybrid layers, is largely responsible for processing hierarchical structure. The hybrid advantage also vanishes on repeated $n$-grams. Overall, while attention suffices for tokens involving copying and purely syntactic information, hybrids show an advantage on more semantically conditioned predictions, likely because these are aided by a strong representation of the discourse state. We conclude with discussion and proof-of-concept experiments showing how these findings could refine design principles and pretraining evaluations for hybrid architectures.
Chat is not available.
Successful Page Load