Evaluability as a Fidelity Dimension: A Case Study in Proposal-Based AI Assistance
Abstract
In this work, we study the problem of modelling user responses to AI proposals in a proposal-based sequential assistance task, which is an abstraction of common human-AI interaction scenarios such as AI-assisted coding. In such settings, the AI often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. Based on the hypothesis that simulating user response in this case may require representing both the value of the proposed outcome and the user's ability to evaluate the proposal, we examine: if a user’s ability to evaluate an AI proposal has a systematic impact on their response, then how should this factor be represented in the user simulator, and what consequences would such an assumption have? As a first attempt, we formulate proposal-based sequential assistance and derive a minimal mechanistic accept-reject model grounded in information-theoretic bounded rationality, explicitly representing both user preference and proposal evaluation burden. We then study the behavioural and proposal planning consequences of this model by deriving acceptance and information frontiers, showing that proposals likely to be accepted need not yield responses that are most informative about the latent user. We further empirically examine the downstream consequences of this model in a controlled graph-based proposal planning setting, which shows that planners with different awareness and representations of user evaluability produce distinct proposal policies and substantially different task outcomes for the same simulated users. Overall, our results characterise how an evaluability assumption propagates through closed-loop assistance and motivate evaluability as a candidate fidelity dimension for proposal-based user simulation.