Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
Abstract
Humans solve complex problems by flexibly shifting among reasoning modes, often without explicit deliberation: they plan, execute, revise intermediate goals, resolve ambiguity through associative judgment, and apply formal procedures to well-specified subproblems. Current LLM agents lack this flexibility, as their scaffolds hard-code such reasoning decisions in advance through fixed inference patterns. These scaffolds are effective when their prescribed structure matches the task, but brittle when solving the task requires adapting the structure of reasoning itself. We introduce Deep Reasoning -- an inference-time approach for constructing task-specific scaffolds through structured meta-reasoning. Deep Reasoning uses a formal language that represents meta-reasoning as executable decompositions over associative inference, formal computation, and recursive subproblem solving, enabling decomposition principles to be encoded as in-context examples that guide test-time scaffold construction. We instantiate this approach in a general-purpose agent (DOLORES) that distributes complex tasks across smaller, more controlled reasoning threads while preserving dependencies among subproblems. We evaluate DOLORES against state-of-the-art scaffolding methods across four hard benchmarks: grounded multi-hop reasoning, synthetic long-chain question answering, long-context aggregation, and deep research-style information seeking. DOLORES outperforms all evaluated scaffolds across four benchmarks, three model sizes, and two model families, improving over the strongest evaluated scaffold baseline by 24.8% on average, including methods tailored to individual benchmark families. Trace and token analyses suggest that while baseline scaffolds fail by overloading individual LLM calls, DOLORES succeeds by distributing cognition across structured, lower-load reasoning threads, thereby reducing premature termination and hallucination. This advantage can even bridge the scaling gap, with an 8B version surpassing all evaluated 32B baselines from the same family in more than half the settings. These results point toward future agentic systems that treat scaffolding as adaptive reasoning, constructing the structure each task requires just-in-time.