Pretraining Amortised Samplers via Unsupervised Skill Discovery
Inhyuck Song ⋅ Jinkyoo Park ⋅ Esmeralda S Whitammer ⋅ Sanghyeok Choi
Abstract
Sequential amortised samplers draw a sample through a sequence of actions and are trained to match a target distribution $\pi(x)\propto R(x)$. We propose a pretraining strategy for such samplers inspired by unsupervised skill discovery: a skill-conditioned sampler is learnt so that each skill covers a distinct region of the sample while their mixture recovers a reference distribution that is easy to sample. We cast this as optimal transport between a uniform skill prior and the reference measure with a cost derived from the Wasserstein dependency measure. The pretrained sampler is then adapted to a downstream target with a skill-conditional trajectory balance loss. Experiments show that discovered skills are geometrically separated and that adaptation from them improves mode coverage and efficiency over baselines.
Chat is not available.
Successful Page Load