Sparse Value of Continued Inference: Learning Adaptive LLM Computation
Ansh Kumar
Abstract
Often, multi-stage LLM workflows can pay for additional inference even when earlier computation is sufficient. We study whether information produced during a first model call can be used to predict when, or if, a second call is worth its cost. On each two-hop MuSiQue question, a first Qwen3.5-35B-A3B call produces a structured reasoning record containing two linked subquestions, a proposed answer, and supporting citations. We store this record, together with deterministic consistency checks, as an external "intermediate state." A ridge-regression controller combines information from this state with summaries of the model's hidden activations to predict the expected delta (improvement) from running a second model call; if the predicted improvement exceeds a fixed threshold, the system continues, and otherwise returns the first-call answer. The second call, when used, receives the original evidence together with the structured first-call record and may revise the answer and citations. On 128 MuSiQue cases (not used to fit or tune the controller), continuation value is sparse: always running the second call improves performance on only 10 cases, harms one, and leaves 117 unchanged. The adaptive policy achieves a composite answer-and-citation score of $0.8239$, compared with $0.8279$ when the second call is always run. It reduces total input/output token positions processed by $34.6\\%$ and measured warm execution time by $24.0\\%$, while identifying 9 of the 10 cases that benefit from continuation. The structured first call alone also outperforms a direct one-call Qwen baseline ($0.8063$ versus $0.7293$), and the adaptive policy outperforms a router that makes its decision from the original question alone ($0.8239$ versus $0.8091$). These results suggest that the intermediate state produced by ongoing inference contains useful information about the marginal value of additional computation, enabling substantial savings with little observed loss in task quality, and support state-dependent selective continuation as a practical form of adaptive LLM computation.
Chat is not available.
Successful Page Load