Varying-Width Vision-Transformers for Efficient Vision–Language Models
Abstract
Research on efficient neural networks has primarily focused on overall model scale rather than the optimal distribution of a fixed parameter budget across layers. Recent work on transformer-based language models shows that non-uniform layer-width allocation significantly impacts learning dynamics. We bring this question to vision–language models (VLMs), whose vision encoder is a Vision Transformer (ViT) that conventionally uses the same width at every layer, leaving open how a fixed budget should instead be distributed across its depth. We evaluate four non-uniform width profiles for the ViT encoder across three parameter budgets on spatial reasoning benchmarks. Under strict parameter and compute parity with the uniform baseline, non-uniform allocations, most notably a diamond-shaped profile that expands its middle layers, consistently achieve lower loss and higher task accuracy. Mechanistically, these architectures retain more decodable spatial information in intermediate representations and transfer faster and more accurately to novel compositional tasks. Our findings provide practical architectural guidelines for maximizing the representational efficiency of vision encoders under fixed resource constraints.