LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
Abstract
Although large language models (LLMs) have tremendous utility, trustworthiness is still a chief concern: models often generate incorrect information with high confidence. While contextual information can help guide generation, identifying when a query would benefit from retrieved context and assessing the effectiveness of that context remains challenging. We introduce LLM Microscope, a framework for auditing LLM outputs directly using model internals. Our key insight is that, while LLMs are poorly calibrated when asked to self-assess via prompting, their internal activations already encode reliable signals about both answer correctness and context utility. We formalize contextual relative utility to distinguish amongst correct, incorrect, and irrelevant contexts. Experiments on six different models reveal that a simple classifier trained on intermediate layer activations of the first output token can predict output correctness with about 75% accuracy, enabling early auditing before full generation. Our model-internals-based metric significantly outperforms prompting baselines at distinguishing between correct and incorrect context, guarding against inaccuracies introduced by polluted context. These findings establish practical tools for better understanding the underlying decision-making processes of LLMs.