GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
Adam Byerly ⋅ Daniel Khashabi
Abstract
Large language models (LLMs) exhibit pronounced position bias in long-context retrieval, systematically prioritizing information location over relevance. Existing mitigations modify positional encodings or attention internals, but in the common output-only deployment setting these are inapplicable. We introduce GOLD PANNING, a Bayesian framework for inference-time active search that (i) reorders documents to concentrate high-belief items in highly diagnostic positions (signal anchoring) and (ii) updates relevance beliefs from model outputs. Unlike active learning, which prioritizes uncertainty reduction, GOLD PANNING exploits anchoring---once flagged, keep it in sight---to preserve weak cues. A greedy assignment derived from the model's diagnosticity profile provably identifies a target among $N$ documents with high probability in $O(\log N)$ rounds. Across open-weight and closed-source models on multi-document QA, GOLD PANNING matches Permutation Self-Consistency's $F_1$ with $30$--$65\%$ fewer queries and remains effective under calibration mismatch, indicating that coarse positional ordering alone drives the gains. These results show that inherent model biases need not be failures, but can serve as exploitable signals.
Chat is not available.
Successful Page Load