Imagine, Don't Narrate: The Generative Bottleneck in World Models of Interaction Dynamics
Daniel Platnick ⋅ Marjan Alirezaie ⋅ Hossein Rahnama
Abstract
Modern AI agents can plan, reflect, reason, and act over long-horizon digital tasks. Even so, they cannot plan to steer a social interaction in real time. Doing so requires anticipating how latent factors like trust and resistance will evolve under the agent's actions, fast enough to plan within a conversational turn. Generative world models approach this by narrating possible futures, but autoregressive text generation is both too slow for real-time planning and fundamentally lossy. Representation-predictive methods can be superior, but are underexplored for interaction. We build LID-Bench, the first controlled testbed with oracle latent states for interaction dynamics, enabling systematic comparison of generative and representation-predictive world models against known dynamics. A bottleneck decomposition on three generative models spanning 117M--350M parameters reveals that all learn interaction dynamics internally, with text-rendering signal losses of 89--100\% observed by horizon $k{=}5$. Our representation-predictive model (Social-JEPA, 125M params, 500K trainable) bypasses text entirely, forecasting latent dynamics $2.8$--$3.9\times$ more accurately while performing rollouts $1{,}059$--$2{,}314\times$ faster---making real-time planning within conversational turn-taking latencies feasible.
Chat is not available.
Successful Page Load