TAPE: A Training-Free Remedy for Temporal Amnesia in Streaming Video Understanding
Mingyu Jeon ⋅ Woojun Jung ⋅ Hyeondong Woo ⋅ Jinkwon Hwang ⋅ Junyeong Kim
Abstract
Streaming video understanding requires access to past events within fixed memory budgets, but whether memory reduction preserves their temporal structure remains unclear. Our controlled temporal-shift evaluation increases query-to-event distance while fixing the target event and question semantics. Accuracy declines across three recent methods spanning KV-cache eviction, selective retrieval, and event-level aggregation. Available matched uncompressed counterparts remain stable, implicating memory reduction rather than relative-time reasoning or additional post-event context. This loss is selective: semantic recall remains stable while temporal precision drops by up to **16%**. We call this **Temporal Amnesia**: remembering *what* happened while losing precision about *when*. We propose **TAPE** (**T**emporal **A**nchoring and **P**roportional **E**ncoding), a simple, training-free remedy for two recurring memory vulnerabilities. *Pinning* preserves at least one anchor per temporal window for coverage; *Reflow* proportionally remaps surviving positions to preserve relative temporal distances. On full RTVU at $\Delta = 30$, TAPE reduces the largest temporal accuracy drop from **8.3%** to **1.3%**, without additional training or parameters.
Chat is not available.
Successful Page Load