When Call-Level Attribution Fails in Iterative LLM Workflows: Representation, Calibration, and Non-Singleton Repairs
Abstract
When an iterative LLM workflow fails, the final output rarely reveals which call introduced the error. Later calls inherit upstream drift, while stochastic calls can diverge even when nothing is wrong. We study three limits on call-level attribution: the representation used to compare states, calibration across a whole trajectory, and the assumption that repairing one call can restore the outcome. To expose these limits, we build a residual scorer that learns normal per-position variation from repeated clean runs and measures each call’s output deviation conditional on its input deviation. We evaluate it by injecting faults into one six-stage workflow on a controlled task, using three open-weight and two frontier model families. State representation matters more than scorer complexity. Across 66 open-weight incidents, replacing lexical distance with a task-aware per-field distance raises exact localization accuracy from 0.24 to 0.47 for the residual scorer and from 0.06 to 0.73 for a simple earliest-threshold-crossing baseline. The residual scorer nevertheless underperforms that baseline (0.47 versus 0.73; paired p < 0.001): it degenerates at the first position and can treat a later sequential fault as inherited drift. Call-level calibration also understates operational false alarms: a 0.04 per-call alarm rate corresponds to alarms on 0.24 of six-call trajectories. Finally, roughly half of valid multi-fault trials require repairing both faults; no single-call repair succeeds. Together, these controlled results show that call-level attribution depends not only on the scorer, but also on task-relevant state representations, trajectory-level error control, and repair sets that allow more than one culprit.