Ensembling Language Models with Sequential Monte Carlo
Abstract
Combining the predictions of multiple models is one of the most effective strategies in machine learning. Thus, it is only natural to investigate the ensembling of language models (LMs). However, because LMs define distributions over strings, a proper ensemble requires aggregating predictions at the string level, which requires a computation that is generally intractable. In this work, we cast LM ensembling as an inference problem and propose the deployment of sequential Monte Carlo (SMC), an appropriate approximate inference scheme. Concretely, we introduce a unified framework for composing multiple LMs into ensemble distributions parameterized by a broad family of aggregation functions. To sample from these distributions, we introduce an SMC algorithm that operates in a shared character space, enabling ensembles of models with mismatching vocabularies and consistent sampling in the limit. We evaluate a range of ensembles across prompt and model combinations for various text generation tasks, finding that consensus-seeking strategies such as the product consistently outperform probability averaging, and that better posterior approximations can yield better ensemble performance.