FASD: Hardware Acceleration for Multi-AI-Agent Discussion
Abstract
Recent development of large language models (LLMs) has empowered multi-agent systems (MAS) through external skills and inter-agent discussion to solve complex tasks. Although such a paradigm improves the capability of LLM systems, it also introduces substantial efficiency overhead. Existing approaches mainly reduce this overhead through software-level optimization. However, these methods overlook how to exploit the graph-level structure of MAS discussion and its suitability for hardware-aware acceleration, thus causing latency, especially when the number of agents, discussion rounds, or concurrent sessions increases. In this paper, we propose FASD, an FPGA-algorithm co-design framework for acceleration towards graph dynamic system agent discussion. We argue that MAS discussion can be formed as a graph-coupled dynamical system and thus agent activation pruning can be optimized through phase-coupled dynamics method. Based on these insights, we designed the FPGA-aware optimization for the problem. FASD mainly consists of the following parts. Firstly, we profile representative MAS workflows and construct a phase profile for dynamic system modeling. Secondly, the hardware-friendly optimization progress, given the profile, is implemented on FPGA. Through extensive experiments on multi-agent discussion workloads, we demonstrate that FASD improves discussion efficiency while preserving task performance, with stronger benefits as the agent graph becomes larger and more complex.