Prysma: Efficient Modality Adaptation for SLO-aware LLM-based Video Question Answering
Abstract
Large language models (LLMs) have demonstrated remarkable performance in video question answering (VQA). To improve responsiveness and answer quality, modern LLM-based VQA services pre-extract multiple LLM-compatible modalities from videos before response generation. However, the performance impact of modality combinations (MCs) and extraction knob settings has barely been studied. Our analysis shows that both factors substantially affect answer quality and Time-to-First-Token (TTFT) latency. Furthermore, the best configuration depends on both video content and query semantics, necessitating adaptive strategies. We propose Prysma, the first modality adaptation framework for LLM-based VQA. Prysma intelligently adapts the modality configurations to optimize answer quality, subject to a Service Level Objective (SLO) on average TTFT. It combines an offline modality orchestrator and an online semantic adapter. The offline orchestrator constructs a video knowledge base to warm-start tuning, employs a latency-aware configuration pruner to reduce search space, and runs a two-stage tuner to optimize both video- and query-level configurations. The online adapter selects configurations in real time based on query semantics and tuning history. Extensive evaluations demonstrate that Prysma improves answer quality by up to 82% over SOTA methods under the same SLO setting.