RSIHub: A Composable Framework for Agent-Evolution Methods
Zimu Wang ⋅ Xiaobo Wang ⋅ Siyu Ye ⋅ Silin Chen ⋅ Yuling Shi ⋅ Ruobing Wang ⋅ Jingjing Zhang ⋅ Chaofan Wang ⋅ Kai WU ⋅ Mengnan Qi ⋅ Kai Cai ⋅ Xiaodong Gu ⋅ Jiaqi Li ⋅ Zilong Zheng
Abstract
Developing a language-model agent is typically an iterative process in which researchers revise prompts, tools, skills, and orchestration after observing its behavior. Agent-evolution methods automate parts of this process by using execution feedback to propose and evaluate persistent revisions. Despite their diversity, implementations of these methods rebuild the same experimental machinery and adopt their own evaluation protocols, making new methods costly to build and difficult to compare under shared conditions. We present RSIHub, an infrastructure that separates the policy of an agent-evolution method from the shared machinery used to execute and compare methods. Within RSIHub, a recipe describes this policy: which candidate to work on, what evidence to use, how to propose a revision, and what to retain. RSIHub supplies shared execution, evaluation, recovery, and lineage services. It keeps the evaluator outside the candidate's mutable surface and ties each result to the exact candidate and evaluation conditions that produced it. For evaluation, we implement recipes for four methods: AHE, Hyperagents, A-Evolve, and GEPA. We evaluate these recipes with MiniSWE and Codex on Terminal Bench 2 and $\tau^3$ Banking under fixed model and evaluator configurations. The largest observed overall pass@1 gain is 13.5 percentage points, while rankings among the evaluated recipes vary across agents and benchmarks, and optimization gains do not always transfer to held-out tasks. RSIHub makes such differences measurable under one protocol. Our code is available at https://anonymous.4open.science/r/rsihub.
Chat is not available.
Successful Page Load