Learning to Audit ML Models with Theory of Mind
Abstract
When auditing machine learning models, human experts sequentially inspect model predictions and explanations, gradually building a mental model of the model’s behavior to uncover errors. This manual process is critical for safe deployment in high-stakes domains but fundamentally limited in scalability. In this work, we propose a novel method to simulate this human auditing behavior for automating the auditing process. We formalize auditing as a sequential decision-making problem, modeled as a Markov decision process. Our state representation summarizes the auditor’s accumulated knowledge of model predictions and feature usage. Building on this formalization, we solve the MDP using reinforcement learning, yielding RLAuditor, which learns to select informative test samples to efficiently uncover model errors. We further connect our formalization to Theory of Mind, showing that our state mirrors how human auditors build mental models: its update rule is equivalent to Bayesian belief revision and it converges to the model’s behavioral pattern at an interpretable rate. We validate our approach on multiple ML models and datasets across different data modalities. In a human subject study, we demonstrate that RLAuditor helps human auditors produce more accurate audit reports in less time.