Contrastive Pretraining Scales Agentic Exploration
Abstract
This paper concerns the problem of scaling agents to large action spaces demanding pretrained world knowledge. What type of representation pretraining would enable policies such as language models to generalize rewards to appropriate actions? Contrastive learning aligns similar inputs and otherwise enforces distancing of representations, while non-contrastive alternatives omit distancing terms and so learn anisotropic representations that collapse distinctive clusters. Though contrastive pretraining has been proven to promote supervised learning efficiency, non-contrastive pretraining nevertheless remains canonical for language agents. Here we examine the extent to which these pretraining types support on-policy finetuning, in which efficiency is coextensive with online performance. Broadly, we argue that contrastive pretraining optimizes efficient exploration of large spaces. We prove that anisotropic representations collapse distinct actions and necessitate complex policies, whereas contrastive pretraining retains distinctions and enables reward transfer through spectral regularization. We validate our conclusions empirically by providing spectral analyses of language data and simulations involving multimodal personalization policies as well as self-improving reasoning agents finetuned online using a policy gradient. Our latter simulations suggest contrastive pretraining for language agents exploring diverse reasoning paths.