Anytime-Valid Partial Identification for Policy Evaluation under Overlap Violations
Abstract
Off-policy evaluation requires the target policy to be supported by the logging policy. This condition fails when an experiment randomizes among finitely many action levels, but the target policy prescribes continuous intermediate actions. In this setting, the policy value is not point identified. Here, we develop a partial identification framework that quantifies the uncertainty in policy evaluation under such overlap violations. We associate the target policy with an identified surrogate policy supported by the logging policy and bound the discrepancy between the surrogate and the target intervention through a sensitivity model. This yields a decomposition of sampling uncertainty and uncertainty due to limited action support. We then combine e-processes with the identification bounds to obtain finite-sample, anytime-valid confidence sequences for the target policy value.