SEMARC: Sequential Modality Acquisition for Efficient Multimodal Inferenc
Abstract
Multimodal temporal models often process every available stream, tying inference cost to availability rather than the evidence needed for each prediction. We introduce SeMARC, a framework for proactive sequential modality acquisition that couples adaptive selection with incremental fusion. Its sequential multimodal bottleneck backbone, SeMA, reuses accumulated multimodal state and predicts after every acquisition, while a post-trained Adaptive Runtime Controller, ARC, uses acquired evidence to select the next available modality or stop. Each decision precedes modality-specific processing, so unselected encoder and fusion branches never execute. Across six datasets and eleven baselines, SeMARC achieves the highest macro-F1 on five datasets, with a 3.2% mean relative improvement and 25.0% fewer GFLOPs on average than the corresponding highest-F1 baselines, while processing 47.7% of available modalities. Across GPU, CPU, and Android, measurements over 13 dataset–platform conditions show average latency and energy reductions of 43.1% and 40.0%, respectively. Under 0–40% test-time missingness, SeMARC remains Pareto non-dominated in 28 of 30 dataset–missingness conditions. These results demonstrate the benefits of coupling adaptive acquisition with reusable sequential fusion for efficient multimodal temporal inference.