QueryStop: Dynamic Stop Signals for Efficient Streaming Inference
Abstract
As language models are increasingly deployed in long-context and streaming workflows, deciding when to stop context processing is central to inference efficiency. While recent research has shown that a subset of attention heads encodes sufficiency signals, these markers are typically detected via static, task-agnostic global features. In this paper, we demonstrate that stop signals are inherently query-relevant and introduce \textsc{QueryStop}, an attention-aggregation method that systematically identifies query-focused attention heads whose hidden features effectively align with the input query. Driven primarily by query-focused heads, this tailored ensemble yields high-fidelity representations that significantly improve sufficiency prediction. Across four benchmarks on eleven datasets, our method improves sufficiency stop prediction by 3.52\% F1 on average. Under comparable token reduction, it raises Sufficiency Reach Rate (SRR), our proposed stopping-adequacy metric, from 87.12\% to 94.73\% and yields a 7.09\% average gain in downstream task accuracy over the prior baseline. Collectively, these results show that \textsc{QueryStop} effectively eliminates unnecessary context processing while simultaneously improving downstream answer quality.