BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces
Abstract
Many decision-support systems require adapting to individual users, yet evaluating this capability remains challenging because user preferences and beliefs are rarely stated explicitly. Instead, they are revealed through sequences of real-world behavior. Existing benchmarks for user understanding often rely on textual personas, constructed preferences, or simulated users, which may not faithfully reflect how people actually decide.We introduce \textsc{BehaviorBench}, a benchmark for modeling user decisions from real-world behavioral traces. Built from public prediction-market and on-chain records, the benchmark reconstructs longitudinal decision histories and evaluates whether models can infer how a particular user believes and acts from prior behavior. It defines two complementary task layers: \emph{Belief prediction}, which infers a user's final revealed stance and confidence, and \emph{Trade prediction}, which predicts the direction and magnitude of individual actions. We evaluate frontier and open-weight generative models under multiple history representations, including direct behavioral history, generated user profiles, and retrieved cross-user evidence. Results show that personalization is not a single capability: models benefit differently from different representations depending on the task, and performance remains limited when reasoning over implicit behavioral signals. \textsc{BehaviorBench} provides an evaluation setting for studying personalized decision modeling grounded in real-world behavioral evidence rather than simulated users alone.