LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
Abstract
Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, cost, and latency. As a result, effective MLLM routing requires more than estimating query difficulty: a router must match the multimodal requirements of the current image-question input with the capabilities of each candidate model. We propose \textsc{LatentRouter}, a router that formulates MLLM routing as counterfactual multimodal utility prediction. Given an image-question query, \textsc{LatentRouter} extracts learned multimodal routing capsules, represents each candidate MLLM with a model capability token, and performs latent communication between these states to estimate how each model would perform if selected. A distributional outcome head predicts model-specific counterfactual quality, and a bounded capsule correction refines close decisions while limiting the influence of noisy routing signals. The resulting utility-based policy supports both performance-oriented and performance-cost routing, and can handle changing candidate pools through shared per-model scoring with availability masking. Experiments on MMR-Bench and VL-RouterBench show that \textsc{LatentRouter} outperforms strong fixed-model, feature-level, and learned-router baselines. Additional analyses show that the gains are strongest on multimodal task groups where model choice depends on visual, layout-sensitive, or reasoning-oriented requirements, and that latent communication is the main contributor to the improvement. The code is available at: https://anonymous.4open.science/r/LatentRouter-8718.