Latent Debate: A Surrogate Framework for Interpreting LLM Thinking towards Binary Decisions
Abstract
Understanding the internal “thinking” process of Large Language Models (LLMs) and the cause of hallucinations remains a key challenge. To this end, we introduce latent debate, a structured surrogate framework for interpreting model outputs on True/False prediction tasks through the lens of internal latent arguments and interactions amongst them. Unlike human debates, latent debate captures the hidden supporting and attacking signals that arise within a model during a single inference. We first present a model- and task-agnostic conceptual framework, and then instantiate it symbolically to approximate the thinking process of LLMs towards binary decisions. Empirical studies demonstrate that our latent debate is a faithful structured surrogate model that has highly consistent predictions with the original LLM, while providing a form of interpretability. We also demonstrate that our latent debate provides a strong baseline for hallucination detection. Specifically, we identify strong correlations between debate patterns and hallucinations, such as a high degree of disagreement in the middle layers of the latent debate surrogate is linked to a higher risk of hallucinations. Our findings suggest that latent debate shows potential to analyze internal signals in LLMs for binary decision settings.