Lightweight, Provenance-Aware Retrieval Gating for Vision–Language Inference on Resource-Constrained Devices
Souparna Chatterjee ⋅ Suchetana Chattopadhyay
Abstract
Vision–language models (VLMs) are too expensive to run on every image of a locally held archive — a dashcam buffer, a fixed-camera store, a field-inspection dataset — on single-GPU, no-cloud-round-trip hardware where such archives actually live, entirely on-device with no external inference calls. Retrieval gating (encode once with a cheap embedding model, run the VLM only on the top-$k$ retrieved candidates) is the standard fix, but a real device budget has several dials, and it is not obvious which are worth tuning once the others are fixed. Instrumenting all of them on a 2,500-image offline archive (Flickr30k + BDD100K) served locally by frozen CLIP, FAISS, and SmolVLM2-256M on one NVIDIA T4, we find a counterintuitive result for on-device precision choices: INT8 quantization — the default “efficiency” lever in most deployment checklists — runs $7.69\times$ slower than FP16 in this batch-size-one, single-model regime, while NF4 delivers the larger memory saving ($35.2%$) at only $1.27\times$ latency cost. This inversion holds because the budget is dominated by a single stage ($c_{\mathrm{vlm}}/c_{\mathrm{emb}} \approx 410$, $c_{\mathrm{vlm}}/c_{\mathrm{query}} \approx 2.3\times10^{4}$), which also means approximate search saves a negligible $0.07,\mathrm{ms}$ against a $1.4,\mathrm{s}$ VLM call, and gate width $k$ is the only dial with a measurable effect on end-to-end cost. A lightweight ($0.26$M-parameter) linear adapter improves retrieval in both domains without closing the gap between general photographs ($R@5 = 0.890$) and a visually homogeneous driving-scene corpus ($R@5 = 0.344$). Because the dominant stage leaves so much headroom, the gate can also emit a structured provenance record — which candidate reached the model and why, versus what the model said — for free, giving fully local, auditable multimodal inference with no cloud dependency for logging or review. Our claims are scoped to one deployment scale and one GPU; the transferable lesson is that on-device quantization and indexing choices should be measured against the stage that actually dominates the budget, not assumed from server-scale intuition.
Chat is not available.
Successful Page Load