Interleaved Latent Thinking and Adaptive Termination for Efficient Reasoning LLMs
Abstract
Emerging latent reasoning paradigms allow Large Language Models to “think” in latent embedding spaces, offering a more efficient alternative to explicit Chain-of-Thought. However, they generally face two critical limitations: First, relying exclusively on latent tokens amplifies uncertainty due to their inherently high entropy, thereby degrading solution correctness. Second, the overthinking problem remains prevalent, and existing approaches typically rely on rigid heuristics for stopping, often resulting in suboptimal termination. To address these challenges, we view efficient reasoning as a learnable control problem over both how to think and when to stop. We formulate this within a Reinforcement Learning framework, training lightweight per-step gates on a frozen LLM. This modulation enables the model to (i) adaptively interleave latent soft-token reasoning with explicit token generation, and (ii) predict cumulative exit probabilities for robust early termination. Experiments on diverse reasoning benchmarks across varying model scales and families demonstrate that our method surpasses both standard CoT and existing latent baselines, achieving accuracy gains of ∼3% while reducing token consumption by up to 40%.