The User (Simulator) is More Realistic than the Agent: Probing Asymmetry in Simulated Coding Sessions
Abstract
Simulated users increasingly stand in for people in the evaluation of multi-turn coding agents. However, we know little about what either the user or the agent understands about the coding tasks they work on together. Task reward and agent behavior are the primary metrics used to study these simulated coding environments. If we could instead measure what simulated users and coding agents understand about their broader task objective, we could ask more fundamental questions about the friction between users and agents during live rollouts. We argue that simulated users are ideal for this objective, given that they are more easily probed than human users. However, much like how human user studies require careful design, the probing mechanism of simulated users and coding agents must be designed and validated with care. To that end we introduce fork-and-discard probing. At a chosen step we copy the subject's frozen state onto a disposable branch, ask it a question to gauge its understanding of the session, and discard the branch, so the interrogation never disturbs the rollout and the same moment can be re-queried as often as we like. Probing ten personas, three coding-agent configurations, and five Terminal-Bench-style tasks, we first find that a probe's answer is noisy and shifts with how the question is framed, so each probe must be characterized before it can be trusted. Characterized this way, our probes reveal that coding agents are poor judges of their own sessions, overconfident about both their success and their pace, while simulated users track true reward and remaining work more closely. This asymmetry in understanding between users and agents points to a clear target for future work: coding agents that judge their own progress as honestly as the users they serve, and act on it before friction sets in.