Learning Evidence Highlighting for Frozen LLMs
Abstract
Large Language Models (LLMs) can reason well over focused inputs, yet often miss decisive evidence buried in long, noisy contexts. We introduce HiLight, an Evidence Emphasis framework that turns evidence selection into a learned, non-destructive input-side control problem for frozen LLMs. Rather than retrieving, pruning, compressing, or rewriting the input, HiLight trains a lightweight Emphasis Actor to insert minimal highlight tags around pivotal spans while preserving the original context. A frozen Solver then reasons over the emphasized input. The Actor is trained only from the Solver's downstream task reward, requiring no evidence labels, Solver gradients, logits, or internal activations. This yields a solver-compatible alternative to retrieval-style hard selection, context compression, and instance-level prompt rewriting. Across sequential recommendation and QA, HiLight consistently improves over manual prompting and strong automated prompt-optimization baselines, with gains up to +27.5\% over manual instruction (MI) and +10.8\% over the strongest baseline on Amazon-Beauty. The learned emphasis policy transfers zero-shot to smaller and larger unseen Solvers across model families, including an API-based Solver, and its highlights align with human supporting evidence up to 0.78 F1. These results suggest that evidence selection can be learned as a reusable input-side control mechanism for frozen LLMs.