HALO-VGGT: Heterogeneity-Aware Lightweight Online Compression Allocator for Efficient VGGT
Xueling Wang ⋅ Yiwen Wang ⋅ Siqi Cai ⋅ Chen Zhang ⋅ guanghui He
Abstract
Visual geometry grounded transformers are strong feed-forward backbones for multi-view 3D reconstruction, but their global attention becomes the dominant inference bottleneck as views and tokens increase. Existing acceleration methods mainly focus on designing compression operators, while compression strengths are often assigned by uniform schedules or offline calibration, overlooking the heterogeneous sensitivity of layers and attention heads during inference. We therefore propose HALO-VGGT, a lightweight online compression allocator requiring no training or calibration for efficient VGGT-style reconstruction. HALO uses a sampled query-key margin proxy to estimate compression sensitivity at the layer and head levels, and assigns hierarchical compression ratios accordingly. Instead of introducing a new compression operator, HALO complements pruning, merging, and sparse attention mechanisms by deciding where and how aggressively they should be applied. HALO improves the trade-off between speedup and accuracy across datasets and backbones, achieving up to 4.20$\times$ speedup and further improving existing merging and sparsity methods by up to 1.46$\times$.
Chat is not available.
Successful Page Load