Necessary or Sufficient? Interpreting LLM Explanations With Behavioural Evidence
Urja Pawar ⋅ Rajitha Ramanayake ⋅ Nabeel Kemal ⋅ Ashwin Kandath ⋅ Owen O'Neill ⋅ Houssem Chatbri
Abstract
LLM decision components that can be used within agent workflows are often asked to explain their outputs. Practitioners may use the factors named in these explanations to interpret, debug, or oversee a component. They may treat a named factor as necessary, meaning that changing it would change the output, or sufficient, meaning that retaining it while removing other changeable information would preserve the output. We test whether these interpretations agree with the component's observable behaviour in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenced it. Controlled black-box interventions estimate a necessity score for each factor by measuring how often changing it changes the output, and a sufficiency score by measuring how often retaining it preserves the output. Across eight models from the Claude, GPT, and Gemini families, the mean Spearman correlations between the cited ranking and the necessity and sufficiency scores are $0.349$ and $0.354$ for advisor recommendation, and $0.431$ and $0.580$ for prompt monitoring. Furthermore, an uncited factor scores above the lowest-scoring cited factor in 57.6% of advisor responses under necessity and 58.1% under sufficiency; the corresponding prompt-monitoring rates are 25.8% and 8.9%. The cited top three contain useful information but do not reliably identify the three factors with the strongest measured influence under necessity or sufficiency. Treating the explanation as an observable behavioural report and testing it through interventions provides evidence beyond the output or explanation alone.
Chat is not available.
Successful Page Load