How Much Evidence Should Retrieval-Augmented In-Context Learning Use Under Distribution Shift?
Chen Wang
Abstract
Retrieval-augmented in-context learning faces a budget problem under distribution shift: increasing top-$k$ can improve evidence coverage, but it can also add off-direction, conflicting, or overly long context that a finite reader cannot use. We introduce a theory-guided diagnostic framework based on \emph{effective evidence coverage}: retrieval is useful only while aligned evidence grows faster than reader-facing ambiguity. In a local linear-ICL model, coupled covariate and task-prior shifts create a mixed risk term that aligned retrieval can reduce up to a residual ambiguity term; a conditional transfer result extends the same budget boundary to frozen readers satisfying a prompt-level stability condition. This yields a falsifiable prediction: recall can keep rising after answer accuracy has saturated or declined. Across QA, verification, and NLI tasks, dense top-$k$ sweeps together with oracle and shuffled-context controls expose this recall--accuracy decoupling. The resulting framework separates coverage- and saturation-limited QA from interference-limited verification/NLI, while calibrated selected-context controllers are reported only as operational probes of the diagnosis.
Chat is not available.
Successful Page Load