Amortized Off-Policy Evaluation for Continual LLM Deployment
Abstract
Accurate evaluation is central to RLHF-based post-training and to selecting which LLM to deploy. Standard reward models are trained on fixed datasets of human preferences over a limited set of policies. As RLHF updates the policy, its actions become counterfactual and sometimes out-of-distribution with respect to the reward function training data, and the definition of reward itself may change over time due to the business requirements. This induces a distribution-shift problem which is the main challenge in evaluating a policy with offline data (off-policy evaluation; OPE). However, classical OPE methods are ill-suited for continual deployment settings as they are defined per task, requiring refitting from scratch on new tasks. Tackling this, we propose PFN-OPE, a prior-fitted network that amortizes OPE across a distribution of contextual-bandit tasks. It is pretrained on tasks constructed from a bank of LLM responses scored by reward models. At test time it maps a logged dataset and sampled target responses to a value estimate in a single forward pass, with no per-task fitting. On HelpSteer2 with Qwen and Llama policies, PFN-OPE achieves 2.3 to 54 times lower error than the best of four direct-method baselines across all tested configurations.