Bayesian Test-Time Inference of Task-Aligned Similarity from Weak Interactive Feedback
Weida Liang ⋅ Kenji Kawaguchi
Abstract
Large language model systems repeatedly make test-time decisions by comparing similarity. They retrieve exemplars, rerank candidates, and choose revisions. But similarity is not unique: a candidate can be close lexically, semantically, structurally, or stylistically, while only one of these axes may matter for the current instance. We propose BASIS~(Bayesian Axis Selection via Interactive Signals), a training-free Bayesian framework that treats weak feedback as evidence about this latent task-aligned axis. BASIS maintains a posterior over candidate targets and similarity axes. It updates this posterior from scalar scores, pairwise preferences, or verifier outputs, and then chooses future probes by expected information gain. In a finite noisy similarity model, we show that fixed-axis and fixed-mixture rules can remain suboptimal under axis shift. Across concept recovery, reasoning exemplar selection, response refinement, and adversarial distractors, BASIS consistently outperforms stronger rerankers, static mixtures, and compute-matched best-of-$N$ search under the same candidate pools and feedback budgets. The average gains are moderate because all methods share the same backbone, candidate pools, and feedback signals. Under axis shift, however, the gap widens substantially, showing that latent-axis inference is not fully replaced by better scoring or more search alone.
Chat is not available.
Successful Page Load