LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval
Abstract
Neural models such as YOLO and HuBERT detect local properties such as objects ("car") and emotions ("angry") in video frames and audio clips, with scores in [0, 1]. Lifting these scores to temporal properties over sequences enables applications such as query matching (e.g., "does the speaker eventually sound happy in this audio clip?") and ranked retrieval (e.g., "retrieve top 5 videos with a 10 second scene where a car is detected until a pedestrian is detected"). We formalize this problem of assigning Scores for TempOral Properties (STOPs) over sequences, given noisy score predictors for local properties. We propose LogSTOP, a scoring function that efficiently computes scores for temporal properties represented in Linear Temporal Logic. Empirically, LogSTOP with YOLO and HuBERT outperforms Large Vision / Audio Language Models by at least 16\% on query matching with temporal properties over objects-in-videos and emotions-in-speech, while matching or improving over temporal-logic baselines. On ranked retrieval with temporal properties over objects and actions in videos, LogSTOP with OWLv2 and SlowR50 improves mean average precision over zero-shot text-to-video retrieval baselines by 28\% and 18\% respectively.