Latent-Lens: Visual Perception in Small Language Models Through Latent Communication
Abstract
Bridging independent vision-language models (VLMs) and small language models (SLMs) in a modular multi-agent system typically forces them to exchange information as discrete text, which creates a lossy bottleneck for spatial and fine-grained details. This constraint is especially acute for SLMs, whose limited capacity makes them more reliant on the quality of the input signal. While recent works on latent multi-agent communication have explored latent-space exchange between identical or homogeneous LLMs, propagating rich visual information across model boundaries without text generation has not been studied yet. We propose Latent-Lens, a lightweight codec that bridges the representation gap between a frozen VLM and a frozen SLM directly in latent space, enabling them to communicate purely through hidden states. The codec consists of a domain encoder that projects hidden states from the VLM into the SLM's embedding space, and an attention decoder formalised as a LoRA that teaches the SLM to attend to these visual prefixes, while training only 1.6M parameters (0.2\% of the combined model size). On free-form visual question answering (CLEVR, GQA), Latent-Lens achieves +10.2-28.3pp over a text-communication baseline given the same LoRA budget (+36.4-78.4pp when both agents are frozen); on image captioning (Flickr8k) it improves over the same baseline by +18.5 (+18.9 frozen). Replacing the SLM (TinyLlama-1.1B) or the VLM (InternVL2-1B) with alternative models of different architectures, vocabularies, and hidden dimensions yields consistent gains, confirming generality across heterogeneous model families.