Surfing the Activation Wave: An Information-Theoretic and Geometric Analysis of LLM Hidden States Applied to Speculative Decoding
Abstract
Decoder hidden states from different layers of large language models (LLMs) often encode complementary information. However, existing speculative decoding methods typically rely on a small set of target-model layers chosen using fixed, position-based heuristics. Selecting an optimal subset of layers is inherently a combinatorial problem, making exhaustive evaluation infeasible even for moderately deep models. In this work, we present a principled framework for multi-layer hidden-state selection for speculative decoding. We investigate layer-wise Shannon entropy, aggregate entropy under hidden-state concatenation, and pairwise CKA dynamics across depth, revealing distinct geometric and informational compartments within decoder representations. Motivated by these findings, we propose entropy and CKA-based layer selection algorithms that favor hidden-state sets that are highly informative and minimally redundant. As part of this framework, we introduce token-entropy filtering for hidden states, improving the quality of layer-selection statistics. In EAGLE3-style speculative decoding, our selected sets outperform fixed baseline choices, while providing a clearer explanation of why some defaults can remain competitive. Our results thus provide a principled and practical framework for the analysis and selection of multiple hidden states.