Biomedical AI Agent Efficiency: Benchmarking and Improvement with TraceCoach
Abstract
Biomedical AI agents solve research tasks through iterative reasoning, retrieval, tool use, and code execution, yet evaluation typically emphasizes final-answer accuracy rather than end-to-end resource use. We instrument the open-source Biomni framework and serialize visible trajectories as normalized, step-aligned traces of outcomes, tokens, rounds, model-call and execution time, and tool activity. We benchmark four reasoning-model conditions on Biomni Eval1-100 and CompBioBench-100. Eval1 accuracy ranges from 62\% to 68\% despite a 17-fold difference in mean main-model tokens, while CompBioBench shows wider variation in accuracy (39--78\%), completion, and long-running execution. Trace analysis further shows that costly and incomplete trajectories account for a disproportionate share of resource use and differ in their execution mechanisms across task regimes. We also introduce TraceCoach, which derives reusable efficiency guidance from paired development traces for a frozen target model. TraceCoach shows the clearest resource reductions on Eval1-100; held-out CompBioBench yields favorable but less certain resource-use changes, with higher observed accuracy in both evaluations. These results support trajectory-level resource accounting as a framework for evaluating and improving biomedical AI agent efficiency.