How Fine-Tuning Objectives Shape Layer-Wise Information in LLM Hallucination Detection
Abstract
Under matched conditions, classifier-head fine-tuning (\textsc{cls}) paradigm consistently outperforms the LLM-as-Judge sequence-to-sequence (\textsc{s2s}) on binary hallucination detection. We trace this gap to the representations induced by the two objectives. Static analysis shows that \textsc{cls} learns a task-specific discriminative direction, pushing hidden states into a clear two-cluster geometry with stronger class separation and linearly decodable label information. By contrast, \textsc{s2s} relies on the pretrained language-modeling head to make decisions. This minimizes the pressure to reorganize representations, leaving the hidden states much closer to the pretrained backbone. Our analysis also reveals that the \textsc{cls} signal saturates before the final layer, after which upper layers form a plateau of statistically interchangeable and largely redundant representations. However, endpoint redundancy alone does not show whether these layers are dispensable during training. Further training-dynamics analysis reveals 1) a lower pretrained substrate, 2) a narrow upper-middle \textbf{InfoWindow} where task information emerges, and 3) a later-synchronizing upper plateau that mirrors this signal. This mechanism motivates us to remove layers above the InfoWindow. By truncating roughly 31\% of the network, we obtain early-exit models that match or exceed full-model performance on Qwen3-4B and 8B. The same depth-fractional prescription transfers to Llama-3.2-3B and Mistral-7B-v0.3. We hope this helps guide efficient task-specific LLM design.