TransZoomer: Efficient Transformers for Arbitrary Image Resolutions
Shivani Mall ⋅ João Henriques
Abstract
High-resolution images contain small objects, object parts, textures, and distant details that are often lost by aggressive downsampling. Transzoomer treats high-resolution vision as a retrieval problem over a dense multi-scale descriptor pool rather than dense Transformer token processing. A fixed set of learnable queries retrieves a bounded number of patches at each iteration; a Transformer processes only these tokens, and its outputs recursively update the queries. Crucially, although forward routing is discrete top-$K$, training uses a straight-through estimator (STE): the forward pass is exactly $K$-hot while the backward pass follows a soft similarity distribution. Task supervision therefore reaches the query embeddings, matching projection, and recursive update, making adaptive retrieval trainable end-to-end without an auxiliary routing loss. Experiments on egocentric gaze estimation, object detection, and fine-grained classification show stronger use of increasing resolution while keeping Transformer token-processing cost bounded.
Chat is not available.
Successful Page Load