When Further Realization Is Unnecessary: Amortized Reasoning for Long-Horizon LLM Agents
Abstract
Long-horizon LLM agents solve interactive tasks by repeatedly choosing the next action from the current state. At many steps, a compact decision is sufficient to specify the action, yet the base agent still emits a full realization in its native output format (e.g., code or action text). Once the compact decision determines the executable content of the step, subsequent generation mostly adds formatting and elaboration, yielding redundant realization. This separation is reflected inside the decoder: compact decisions become locally recoverable in middle layer representations before full realizations, and attention over its token span becomes concentrated when the decision is sufficiently supported as the current step output. Building on this separation, we propose MIRA, a training-free inference method for Model-Internal Reasoning Amortization. MIRA amortizes realization through two complementary stages: Consensus-Guided Decision retrieves successful reference steps in the middle layer representation space to propose a compact decision candidate. Attention-Guided Commitment evaluates its token span using token confidence and attention concentration. The candidate is used as the step output only when both signals support it for the current state; otherwise, the agent falls back to the base agent's full realization path. Experiments on long-horizon interactive agents show that MIRA improves task performance while reducing redundant realization, yielding a 47.5\% reduction in output tokens and a 24.7\% reduction in inference time across tasks and agent formats.