Can LLMs explain themselves truthfully with code?
Abstract
Large language models (LLMs) are increasingly used in high-stakes settings where they are expected to justify their predictions, and a common approach to understanding LLM predictions is via self-generated free-text explanations. However, free-text explanations are instance-wise and non-executable, making it hard to evaluate their alignment with model behavior or explain dataset-level behavior in a compact, human-interpretability manner. We study code as an executable alternative to free-text, enabling automatic evaluation and dataset-level explanation, and ask whether LLMs can \emph{consistently} explain their behaviors with code. This automatic verification enables a generation procedure that samples multiple candidate programs and selects the most consistent one. Using three open-source code-capable models and eight tasks with varying levels of linguistic difficulty, we study when LLMs produce consistent code explanations. We find that models can produce highly consistent code explanations in a subset of settings, but strong consistency is not the norm. Consistency is substantially lower for tasks involving more complex linguistic components, where performance is often near random guessing across models. In these harder settings, generated programs are longer but not more structurally complex, and increasingly rely on hard-coded linguistic patterns from the input. These findings suggest two possibilities: either LLMs are not good at producing compact explanations of their behavior, or some behaviors require dense, distributed computation that cannot be explained by compact, human-interpretable programs. Finally, we find that model-internal representations do not reliably predict when code explanations will be consistent, suggesting that LLMs do not know when they can or cannot explain their own behavior with code. Our work highlights code as a compact and verifiable alternative to free-text explanations, while also revealing limitations in using compact, human-understandable explanations to fully understand LLM behavior.