Feeling of Knowing in Large Language Models
Abstract
In real-world deployment, large language models (LLMs) are frequently updated through post-training techniques to maintain up-to-date knowledge. Yet their reliability depends not only on what the LLM knows, but also on whether it knows what it knows—a self-assessment known as the feeling of knowing (FoK). FoK is the signal behind selective generation, retrieval triggering, and model routing; when miscalibrated, it leads systems to overconfidence on unknown questions or to refuse ones they could have answered. As LLMs are updated through posttraining, however, we discover that what an LLM knows can change without a corresponding update to its FoK. This creates a knowledge–FoK desynchronization challenge, where FoK estimators calibrated to the pre-trained LLM become unreliable as the model’s knowledge state evolves. To address this challenge, we propose a two-channel design: a knowledge channel that changes what the LLM knows and a FoK channel that judges whether the current knowledge state supports answering. We train these channels with a three-stage procedure: onpolicy supervised fine-tuning (OP-SFT) updates the knowledge channel from corrected on-policy answers, knowledge-delta FoK training (K∆-FT) trains the FoK channel over intermediate knowledge states, and FoK-guided policy optimization (FGPO) uses the judgment to generate the corresponding answer or explanation. Experiments on three datasets show that explicit FoK alignment substantially improves judgment quality on updated knowledge. The code is available at https: //anonymous.4open.science/status/Know-thyself-code-8D65.