Known Worlds, Unknown Composition: The Value of Action Diversity
Manoj Saravanan
Abstract
Known local models need not determine their correspondence across unobserved switches. We study this obstacle in a two-mode, binary-state model with exact local laws, one unknown sign correspondence, and a known symmetric switching probability $\lambda$, starting in stationarity. For one uniformly hidden boundary between distinct positive laws $p,q$ on a finite alphabet, we prove that boundary laws $r,r'$ have zero limiting observable Kullback--Leibler divergence exactly when $r-r'\in\operatorname{span}\{p-q\}$, with quantitative finite-window bounds. We then classify every fixed finite positive channel library. For sufficiently small $\lambda$ and $0<\delta\le1/4$, the optimal deterministic horizon for error at most $\delta$ under each correspondence is infinite, or has order $\lambda^{-2}\log(1/\delta)$, $\lambda^{-3}\log(1/\delta)$, or $\lambda^{-1}\log(1/\delta)$, according to the channel signatures. At rank one, residuals of one reference channel exactly simulate the observable history law of every causal policy; at rank two, an independent mixture attains the inverse-switch identification order. The proofs combine an equivalent positive hidden Markov representation with observable block separation uniform over different posterior beliefs. Two explicit libraries show that erasing channel labels can slow identification or destroy it. We also give a finite-horizon model-error bound and an $O(\lambda^{-1}\log(2T))$ regret upper bound for an ordered-signature, set-point, additive-reward control subclass. Thus actions and their recorded labels determine which composition information survives hidden timing.
Chat is not available.
Successful Page Load