Latent Space Co-training for Few-Step Text Generation
Abstract
Flow map language models generate text in a few network evaluations by transporting directly between distant noise levels. Existing approaches distil such maps over one-hot token embeddings or the frozen hidden states of a pre-trained language model – in either case the space is fixed before the map is trained. We distil instead inside the latent space of a latent text diffusion model, letting the encoder adapt together with the map, and propose a recipe that keeps this stable. The resulting model, CoFlow, outperforms prior flow-map language models on OpenWebText and GSM8K at small step budgets. In a controlled ablation we show that adapting the space improves perplexity at the smallest step budgets and brings token statistics closer to real text on OpenWebText, and raises problem coverage under repeated sampling on GSM8K.