Would the User Act If We Waited? Human-First Pacing for Real-Time Conversational Agents
Abstract
Real-time proactive assistants must decide when to speak first. Existing systems use multimodal context and conversational pacing to avoid interrupting too early or responding too late, but a timely and non-disruptive intervention can still preempt an action the user would otherwise take. We propose a work-in-progress framework for human-first pacing: estimate what the user would do if proactive AI were held off a little longer, then initiate only after a calibrated opportunity for human choice within a flow-compatible timing region. We separate human-first control, independent initiative, and independent capability. Randomized holdoff identifies short-horizon event-time curves for independent action and explicit help seeking; a stricter independent-action guardrail applies on low-stakes tasks known to be independently feasible. A separate longitudinal audit tests whether repeated exposure changes later independent initiative under a common non-proactive policy, while no-AI probes measure capability. The framework therefore targets natural pacing without treating short-term non-preemption as proof of long-run autonomy preservation.