Intra-Option Fitted Q-Evaluation: Evaluating Hierarchical Policies from Non-Hierarchical Data
Abstract
Off-policy evaluation (OPE) estimates the expected return of a target policy from previously collected data without additional environment interaction. While OPE methods for flat Markovian policies are well studied, little work has addressed evaluating hierarchical policies in which a high-level policy selects temporally extended options that generate primitive actions until a termination condition is met. The primary existing approach applies per-decision importance sampling at the option level, but requires option-annotated trajectories and exhibits variance that grows with the horizon and action dimensionality. For non-hierarchical policies, fitted Q-evaluation (FQE) typically achieves lower error than importance sampling by leveraging the classic Bellman equation for policy evaluation; however, a naive extension of FQE to hierarchical policies results in a biased policy value estimate. In this paper, we introduce Intra-Option Fitted Q-Evaluation (IO-FQE), which estimates a value function in the augmented state-option space; this approach enables OPE of hierarchical policies even when we lack annotations of what option was ran in the data. We develop two continuous-action instantiations of IO-FQE and show on hierarchical continuous control tasks (AntMaze navigation and OGBench Puzzle manipulation) that IO-FQE eliminates the bias incurred by a naive application of FQE to hierarchical policies and substantially lowers mean squared-error compared to option-level importance sampling.