Budgeted Multi-Source Counterfactual Annotation for Off-Policy Evaluation
Biao Xiang ⋅ Ali Eshragh ⋅ Yuexing Li ⋅ Kai Wang
Abstract
Off-policy evaluation (OPE) estimates the value of a target policy from logged data, but limited behavior-policy coverage can force high-variance reweighting or reward-model extrapolation. Counterfactual annotations can add evidence about unobserved actions, yet practical sources, including domain experts and LLMs, may be costly, biased, or noisy. We study budgeted acquisition of such annotations for contextual-bandit OPE. Given source-specific costs and error profiles, we formulate an integer allocation problem over context-action pairs and annotation sources to minimize the component of estimator variance that depends on the annotation plan. We characterize when annotations are valuable through a first-annotation threshold and local annotation-value regimes. For the coupled multi-source problem, we develop a majorization-minimization algorithm with dynamic-programming subroutines that monotonically improves the objective. Experiments in synthetic clinical and LLM-annotated education bandits show that our proposed allocation method reduces the optimized variance component by $20.58$% and $10.77$%, respectively, relative to no-annotation baseline.
Chat is not available.
Successful Page Load