Necessary or Sufficient? Evaluating Explanations from Enterprise LLM Decision Systems with Behavioural Interventions
Urja Pawar ⋅ Rajitha Ramanayake ⋅ Nabeel Kemal ⋅ Ashwin Kandath ⋅ Owen O'Neill ⋅ Guillaume Bourgeon ⋅ Houssem Chatbri
Abstract
LLMs increasingly support enterprise workflows as recommenders or judges and are asked to explain their outputs. We evaluate LLM self-reported ranked explanations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. For each output, models report the top three factors that most influenced it. We use controlled black-box interventions to test two interpretations of these explanations: whether the models select influential factors and whether they rank them in order of influence. Necessity scores measure whether changing a factor changes the output, while sufficiency scores measure whether retaining that factor without other changeable information preserves the output. Across eight models from the Claude, GPT, and Gemini families, an uncited factor outscored the weakest cited factor in about $58\%$ of advisor responses under both necessity and sufficiency, compared with $26\%$ under necessity and $9\%$ under sufficiency for prompt monitoring. The cited ranks were generally positively correlated with the intervention scores, but this varied substantially by use case, model, and criterion. This framework evaluates whether LLM self-reported explanations agree with observable decision behaviour and can support human oversight of enterprise LLM decision systems.
Chat is not available.
Successful Page Load