Lemon: Evidence-Risk-Aware Adaptive Organization for Long-Horizon LLM Agents
Abstract
Large language model agents now combine tools, long context, memory, and multi-step workflows, yet they still fail when execution loses track of the evidence required for a reliable answer. Common failure modes include missing support, unresolved contradictions, and evidence lost under context pressure. We present Lemon, an evidence-risk-aware framework that makes such risks an explicit control variable. At the intra-instance level, a single state-conditioned controller jointly schedules reasoning, tool invocation, worker expansion, verification, context compression, and memory operations. A recoverable evidence substrate pairs anchored context compression with reusable semantic memory, preserving raw tool outputs while storing transferable process fragments. We further extend Lemon to an inter-instance setting in which heterogeneous personalized agents expose expertise, exchange evidence-bearing proposals, critique unsupported claims, and internalize useful collaboration artifacts. Empirically, Lemon reaches 91.36\% accuracy on GAIA and 80\% on xbench-DeepSearch. In an open-source reproducibility comparison on GAIA, Lemon uses 49.9\% to 88.4\% fewer tokens per task than three top-ranked agent baselines.