Asking Is Not Observing: The Elicitation Gap in Chronic-Care Agents
Abstract
Longitudinal chronic-care agents repeatedly ask patients about symptoms, medication use, and changes in daily status, yet they are often evaluated by whether they cover a predefined set of required questions, which assumes the clinical record to be effectively fixed once the right content is queried. That assumption can fail: in interactive follow-up, how an agent asks can change what clinically relevant information becomes observable. We study the problem in a controlled 14-day heart-failure follow-up simulator with privileged latent clinical state, time-varying interaction state, and policy-independent replay. Across 1,614 matched patient-days that hold clinical state, requested content, interaction state and turn count fixed, 996 produce different clinical records; among divergent pairs the plain scripted wording obtains the more complete record in 69.2% (95% CI 66.2–72.0%), with substantial heterogeneity across interaction-state cells. The effect of elicitation strategy also varies with that state: in a preregistered contrast, batching the required items costs disengaged patients 26.7 percentage points (95% CI 19.5–33.9) more of the true record than engaged ones. A policy trained for required-item coverage exposes a second failure mode: its conversational register changes with interaction state while its elicitation structure stays nearly fixed, against a state-aware comparator that moves from 3.30 to 1.24 required items per utterance. Near-perfect coverage (0.998–0.999) co-exists with a third of the true findings in the disengaged cell never entering the record, and re-scoring frozen trajectories on values received separates policies the coverage objective rewards identically. These results motivate evaluating and optimising longitudinal care agents for information arrival rather than question completion.