Spectral Representations for Provably Robust Offline Meta-Reinforcement Learning from Bagged Rewards
Yashas Vaidya ⋅ Bo Dai
Abstract
Context-based offline meta-reinforcement learning (COMRL) aims to learn adaptable policies entirely from static datasets by inferring latent task representations from small contexts. Most COMRL methods rely on immediate, per-step reward signals. This can be generalized by aggregating rewards into ``bags'' and being delivered after a sequence of actions (where bag size $n=1$ is the per-step case). We formally prove that under Bagged Rewards, existing per-transition COMRL methods suffer an information-theoretic collapse ($\mathcal{O}(1/n)$), causing their task inference to degrade to behavior cloning. To solve this, we introduce SpectralMeta, an algorithm that fully decouples dynamics representation from task inference. By pre-training an energy-based spectral representation on reward-free transitions, we recast task identification as a closed-form Bayesian linear regression over bagged rewards. SpectralMeta comes with finite-sample bounds on reward recovery and policy sub-optimality. In continuous-control meta-RL benchmarks, SpectralMeta maintains robust task inference and competitive returns even when rewards are reduced to a single episode-level scalar, a regime where existing baselines fail to recover task-relevant structure.
Chat is not available.
Successful Page Load