From Likelihood Convergence to Parameter Convergence in POMDPs
Abstract
Likelihood-based methods are widely used to learn the parameters of partially observable Markov decision processes (POMDPs), but their finite-sample behavior is not well understood. We establish a different kind of finite-sample guarantee. For any candidate POMDP parameter, we bound its parameter recovery error in terms of its likelihood gap to the ground truth, the structural conditioning of the observation and transition kernels, and the behavior policy used to collect data. The bound applies in both the undercomplete and overcomplete observation regimes, with the overcomplete case following naturally from the undercomplete analysis. Because the result applies to any sufficiently good candidate, it yields finite-sample guarantees for the maximum-likelihood estimator, OMLE-style algorithms, and any procedure that produces a parameter with controlled likelihood gap. This is particularly useful for downstream tasks such as infinite-horizon planning with state-dependent rewards, where accurate recovery of the kernels themselves is required to obtain near-optimal policies. The bound separates the contribution of the behavior policy from that of the model, making explicit how data-collection choices, including action coverage and mixing rate, accelerate parameter learning.