Execution State Attribution: From Provenance Records to Causally Sufficient State in Agent Runs
Aditya Girish ⋅ T D Skand
Abstract
Training data attribution indicates the examples responsible for a response and behavior of a model, and here, we seek commensurate attribution for an agent’s execution – which subsets of an agent’s runtime state are sufficient to explain its observed behavior? While the current techniques of causal attribution focus on necessity of individual execution steps, step-wise necessity does not compose to set-wise explanations.We define \textbf{execution state attribution} as recovery of the minimal subsets of runtime state that reproduce a behavior under resampling of their complement. Formally, a set $S$ is $\gamma$-sufficient when $$ \Pr\left[\beta \mid S=s^{\mathrm{obs}},\operatorname{do}(E\setminus S\sim q_{E\setminus S})\right]\geq\gamma, $$ and our aim is to identify the antichain of minimal $\gamma$-sufficient sets. For a synthetic benchmark, we demonstrate that necessity fails in two simple dependency patterns: under conjunction, individually necessary elements lead to a pair of sufficient elements that cannot be captured by any step-wise score; under redundancy, individually sufficient causes are suppressed and can be misranked. Sufficiency is able to recover the planted explanation at various threshold values. We claim that the estimation of execution-state attribution is a better target for auditing, incident response, and liability, and we point to several technical difficulties for future benchmarks.
Chat is not available.
Successful Page Load