Urgency-Aware Autoregressive VLMs for Unanticipated Healthcare Occurrences
Abstract
The sequential nature of autoregressive decoding in large vision-language models creates a fundamental trade-off between latency and clinical responsiveness when managing asynchronous signals. When unexpected healthcare occurrences arise during an ongoing VLM response, integrating unanticipated events---such as sudden device alerts or urgent caregiver messages---conventionally requires either injecting the new data directly into the active response stream or deferring it until generation ends. Neither approach is ideal in dynamic environments that demand urgent intervention without repeatedly disrupting routine continuations. We present ReflexMind, a training-free framework for online, urgency-aware occurrence handling that runs an Urgency Evaluator (UE) concurrently with the Primary Generator (PG). By leveraging the rotational invariance of Rotary Position Embedding (RoPE), ReflexMind enables a dual-view decoding process that shares a single KV cache, allowing it to assess incoming multimodal occurrences while preserving the state of the ongoing response. This design supports structured decisions, enabling ReflexMind to escalate critical alerts immediately while handling non-urgent occurrences without interrupting the ongoing response. Across three benchmarks spanning multimodal occurrences, our method generally improves occurrence triage, urgency recognition, and intervention selection for VLM backbones at different scales, while using fewer post-occurrence tokens in most settings and adding modest runtime and memory overhead under concurrent occurrences.