Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations
Abstract
When an agent commits to a consequential action after dozens of turns, an overseer needs to know which earlier turns that action actually rests on. Attribution is the natural primitive for answering this, but existing context attribution methods process the full transcript in a single pass: they recover surface-level dependencies while missing the layered structure of real agent traces, in which an assistant turn is itself a derived artifact that condenses, filters, or transforms earlier context. We introduce multi-turn context attribution (MTCA): given a target span in a model response, the task of tracing attribution backward across turns to identify not only which prior turns were directly relevant, but also how those turns themselves depended on earlier context. We propose Tokengeist, an attribution-method-agnostic and scalable framework that recovers full dependency paths by casting attribution as a recursive traversal of a directed acyclic graph (DAG) over conversation turns. We will release MTCABench, a benchmark of 3,845 target spans across 665 multi-turn conversations, (688 ConFETTI, 3,157 TauBench) annotated with gold provenance graphs reaching depths of up to 14, across four dependency types. Across four open-weight models, flat attribution methods fail to recover multi-hop dependencies, with source recall falling as low as 20%, while Tokengeist reaches 90%. Tokengeist does not adjudicate whether a traced claim is correct---it localizes the evidence a downstream verifier or human reviewer must inspect, cutting that material from the whole transcript to a traced path. Our results reveal a systematic failure mode of single-pass attribution---which we term provenance collapse, where attribution terminates at the nearest assistant restatement instead of the originating event---and motivate attribution methods that reason recursively across turns.