Firefly: Illuminating Verified Real-world Tool Call Data Generation
Abstract
Training robust tool-calling agents requires large-scale trajectory data that is both realistic and verifiable, yet existing datasets still rely heavily on costly human annotation or synthetic pipelines with weak correctness guarantees. In this work, we introduce \projectname, a large-scale synthetic and verifiable dataset for tool-calling agents built on real-world MCP servers. Unlike prior pipelines that first generate tasks and then attempt to solve them, \projectname inverts this process: a strong LLM first self-explores real MCP servers and records reachable execution states, after which natural-language tasks are synthesized backward from observed outcomes. This back-chaining design guarantees task reachability and label correctness by construction. To support scalable offline training and evaluation, we further construct a retrieval-augmented simulator from explored trajectories, enabling reproducible tool execution without relying on live MCP servers. Each instance includes a task description, tool schemas, ground-truth trajectories, ground-truth answers, and tool-call responses for offline replay, and is automatically checked for determinism and semantic consistency. \projectname aggregates thousands of verifiable tasks across diverse real-world MCP servers and supports scalable agent learning without human annotation. Experiments show that models trained on \projectname achieve significant improvements on tau2-Bench and BFCL, demonstrating the effectiveness of realistic, verifiable synthetic data for training generalizable tool-calling agents.