Markowitz Regression Oracles: A Mean–Variance Reduction of Reward–Inference Tradeoffs in Bandits
Abstract
At each round of an adaptive experiment, the learner observes a context, assigns one treatment, and sees only that treatment's potential outcome. Reward seeking can therefore suppress the counterfactual observations needed to support a concurrent claim on effect (significance or power). Augmented inverse-propensity weighting-based estimation of average treatment effects (ATE) remain centered under adaptive logging, but its precision depends on conditional outcome variance and on errors in the learned nuisance model, both of which may be magnified by small assignment probabilities. We make this finite-sample dependence learnable rather than assuming that nuisance estimates eventually become accurate: a Markowitz Regression Oracle learns conditional means and second moments, while Neyman Entropic Descent keeps probability on comparisons whose variance geometry remains unresolved. Mixed with inverse-gap weighting, the method attains simultaneous sublinear reward and strong Neyman regret. The same variance certificate governs the clock of an anytime-valid AIPW confidence sequence obtained by inverting an e-process, relaxing a key prior assumption of asymptotic nuisance consistency and instead replacing it by a learnability rate of a regression oracle. On three nonlinear heteroskedastic semi-synthetic tasks, the resulting policy is point-estimate nondominated in reward versus strong Neyman regret and forms the frontier knee on two clinical datasets.