How Should We Train Tractable Proxies for Language Models?
Abstract
Probabilistic inference is increasingly used to guide language models (LMs) to generate text that satisfies logical and semantic constraints. Existing methods train a tractable generative proxy for the LM, typically through maximum likelihood estimation (MLE). Motivated by the success of reinforcement learning in optimizing LMs for downstream objectives, we ask whether tractable proxies can likewise benefit from training objectives that reflect their downstream role. We introduce a variational reverse-KL (RKL) objective as an alternative to MLE for training hidden Markov model (HMM) proxies to approximate the LM’s distribution. Compared with MLE training, RKL better reflects the LM's linguistic behavior: for semantic control, it gives more calibrated signals in the LM's token-probability space; for logical control, it better captures the LM’s relative advantage on concept sets with higher pairwise mutual information. Across both settings, RKL-trained HMMs provide better inference-time signals, improving generation fluency and downstream performance. In decoding, we also add the HMM’s next-token probabilities, focusing guidance on regions where the RKL-trained HMM concentrates probability mass and provides more reliable constraint corrections.