Introspection Tools Help LLMs Understand and Control Themselves
Abstract
Large language models (LLMs) exhibit increasingly strong reasoning capabilities, yet they remain limited in their ability to understand, explain, or regulate their own states and behavior. While interpretability methods have advanced our ability to expose and manipulate their internal mechanisms, they are designed for human analysts rather than for models themselves. In this position paper, we argue for equipping LLMs with tools that expose measurable internal signals (e.g., token probabilities and internal activations) and enable self-regulation. We refer to these as introspection tools, drawing inspiration from humans' access to their mental and physical states, which allow for effective mental and bodily control in diverse real-world contexts. We present the design space for introspection tools, and two proof-of-concept examples illustrating how such tools are feasible and can support uncertainty quantification and misinformation mitigation. Beyond manually designed tools, we highlight open research questions around the automatic development, refinement, and organization of introspection tools. Together, these directions move toward LLMs that are more self-aware and reliable.