JARVIS-Bench: Benchmarking Personal Intelligence Agents on Long-Horizon Real-User Daily Traces
Abstract
Personal intelligence agents aim to support users across daily-life tasks by learning their preferences, routines, needs, and emotional states from long-term interaction. They are expected to anticipate user needs, provide proactive assistance, and infer preferences that users have never explicitly articulated. However, existing benchmarks for LLM personalization lack real user data spanning extended periods and diverse daily-life domains, and typically reduce personalization to simplified settings such as recommendation or fact retrieval. We introduce JARVIS-Bench, a new benchmark for evaluating personal intelligence agents, curated from real users’ daily traces and augmented with synthetic trajectory expansions grounded in implicit-preference insights from surveys and structured personas. JARVIS-Bench is built from participants who contributed 28 to 44 days of multimodal self-tracking logs, structured personas, and a 100-question implicit-preference survey spanning eight domains. It supports four downstream task surfaces: implicit preference reasoning, proactive-help prediction, emotion-aware modeling, and cross-domain transfer, all evaluated through a unified event iterator and long-horizon user-memory pipeline. We benchmark 28 LLMs from 10 model families, including Gemini, GPT, Qwen, Gemma, and Claude, across a five-condition input ladder ranging from no memory to lifelong memory. Results show that memory and model scaling improve personalization, yet current frontier models remain far from achieving robust real-world personal intelligence.