Conservative Continuous-Time Treatment Optimization
Abstract
We propose a conservative continuous-time stochastic control framework for treatment optimization from irregularly sampled patient trajectories. We model the unknown patient dynamics as a controlled stochastic differential equation where treatment is a continuous-time control input. Naive model-based optimization can exploit model errors and propose out-of-distribution controls that appear optimal under the learned model but perform poorly under the true dynamics. To mitigate such extrapolation, we introduce a consistent, signature kernel MMD regularizer on path space, which penalizes controls whose predicted trajectories deviate from observed ones. The resulting conservative objective minimizes a tractable upper bound on the true cost. Experiments on pharmacometric simulators show improved robustness and performance compared to non-conservative baselines.