SIQ2-Long: Evaluating Evidence Retrieval for Long-Horizon Social Reasoning
Abstract
AI systems are increasingly integrated into social environments, such as collaborative workspaces, online communities, and embodied wearable systems. To that end, vision-language models have made significant progress in social reasoning capabilities, with open-source VLMs exceeding 70\% accuracy on the social intelligence benchmark Social-IQ 2.0 \citep{wilf2023siq2}. However, two interacting limitations persist, hampering their reliable deployment outside of well-defined benchmarks. First, VLMs struggle to reason reliably over the time horizons of realistic social interactions, forcing a VLM-backed system to find localized evidence before reasoning over it. In parallel, vision-language encoders are typically trained to align visually-descriptive text with visual content, whereas social questions often require inferential alignment between an observation and its cause (“Why is person X frustrated?”). This call into question the VLM's current ability to serve as a faithful retriever of evidence within long-horizon social reasoning contexts. In this work, we first establish the need for evidence retrieval in long-horizon contexts, evaluating five frontier VLMs on progressively longer video inputs. We show that QA accuracy drops by up to 15.6 percentage points as input size increases from 1 minute to 5–10 minutes, with some models refusing longer inputs altogether. We then present SIQ2-Long, an adaptation of the Social-IQ 2.0 benchmark into long form, enabling both video QA and evidence retrieval testing. We show that, given a multiple-choice question, frozen video-language encoders can identify the right video but fail to distinguish between 1-minute clips within the same video, recovering the correct one-minute “oracle” segment at below chance. Contrastively trained retrieval heads raise top-1 accuracy by 21.7 points, but further analysis suggests that what is learned is more structural than semantic -- which clips tend (or tend not) to be evidence, rather than which clip answers the question. Despite large gains in retrieval accuracy, downstream QA does not budge: decoder-only VLMs answer similarly with randomly selected within-video context, while oracle evidence provides a consistent, if modest, advantage. Our results reflect the two limitations: long contexts necessitate evidence retrieval, yet systems remain capacity-bound even when the relevant evidence is provided, and retrieval itself suffers from over-reliance on descriptive signals. By disentangling where these failures arise, our findings provide a clearer roadmap for building long-horizon social reasoning systems that retrieve and reason over evidence more effectively. SIQ2-Long and code are available at \url{https://github.com/tomcohen13/siq2long}.