Elicitation Matters: How Prompts and Query Protocols Shape LLM Surrogates under Sparse Observations
Abstract
Large language models are increasingly used as surrogates for scientific optimization, where uncertainty helps determine which candidate solutions should receive costly experimental evaluation. Holding the observations fixed, we show that predictions and uncertainty are not determined by the data alone, but change systematically with semantic prompt content and query protocol. We introduce an uncertainty-alignment criterion measuring whether model uncertainty tracks residual ambiguity among sample-consistent functions. Across controlled inference and Bayesian optimization tasks, structural prompts act as effective priors: correct information improves predictions, while incorrect information systematically misleads the surrogate. POINTWISE is generally more observation-faithful and ambiguity-sensitive, whereas JOINT is more internally coherent but less constraint-faithful. These differences change acquisition decisions and regret, showing that elicitation protocol is part of the surrogate specification, not a formatting detail.