Toward Efficient Reasoning of Large Language Models via Latent Concept-Pyramid Modeling
Abstract
Large language models (LLMs) perform reasoning via lengthy token-by-token generation, incurring substantial inference cost. While recent methods compress this process by enabling LLMs to reason in a latent space, they still rely on next-vector generation for sequential logic and on flat representations. We therefore present Latent Concept-Pyramid Modeling (\emph{L}CP), a new paradigm that reformulates LLM reasoning as hierarchical next-level concept generation in a coarse-to-fine manner. Specifically, \emph{L}CP generates the pyramid level by level: a single apex concept encodes the most abstract semantics, and each subsequent level refines its predecessor into finer-grained concepts that together capture distinct CoT segments at one granularity. Such next-level concept prediction approximates the reasoning trace in a manner analogous to advancing from a skeletal outline, through broad structural forms, down to local details. Experimental evaluations on zero-shot mathematical and coding reasoning benchmarks demonstrate that Qwen2.5 and Qwen3 models fine-tuned with \emph{L}CP achieve significant reductions in both token cost and inference latency while maintaining competitive accuracy. \emph{L}CP further showcases zero-shot generalization ability across different tasks. Moreover, the pyramid's hierarchical, multi-grained representational capacity enables faithful reconstruction of the original CoT, endowing \emph{L}CP with interpretability.