Explanation Mechanism Influences Human Reliance on Reinforcement Learning Agents
Abstract
When reinforcement learning agents are deployed in decision-support settings, humans must repeatedly decide whether to adopt or override individual action recommendations. Explainable reinforcement learning supports these decisions, yet it remains unknown whether explanation mechanisms impact user behaviour. We conduct two controlled user studies that isolate explanation mechanism while controlling for agent policy, task, and abstraction level, using a design that independently manipulates action optimality and explanation veracity. Study 1 compares three representative action-level mechanisms across 100 participants and finds that Shapley-based attributions produce conservative reliance sensitive to explanation plausibility; saliency-based explanations indiscriminately increase adoption even for suboptimal actions; and a novel reward-decomposition mechanism, Action Advantage Attribution (AAA), achieves the highest appropriate adoption and is the only condition in which explanation comprehension positively predicts appropriate adoption. Study 2 benchmarks AAA against a no-explanation baseline across 70 participants, showing that it substantially improves detection of appropriate recommendations by 89\% but does not eliminate over-reliance on suboptimal ones. Our results imply that algorithm designers must treat explanation mechanism as a first-order design choice, as mechanisms can influence human reliance on reinforcement learning agents.