Capacity-Constrained Online Convex Optimization with Delayed Feedback
Alexander Ryabchenko ⋅ Idan Attias ⋅ Dan Roy
Abstract
Online learning with delayed feedback typically assumes that the learner can track all pending rounds until their feedback arrives. In practice, tracking resources are finite, and feedback from rounds that cannot be tracked is permanently lost. In this paper, we study delayed online convex optimization (OCO) under a hard capacity constraint, where at most $C$ pending rounds can be tracked at any time. To model delay information, we introduce a semi-clairvoyant model that refines the clairvoyant assumption from prior work: rather than requiring delays to be known at prediction time, the learner observes delay expirations online, consistent with the classical unconstrained delayed setting. Our approach proceeds via a reduction to a novel ``delayed and weighted" OCO problem, using a scheduler that randomizes which rounds are tracked and importance-weights the resulting observations. For the base problem, we propose and analyze Delayed-Weighted FTRL and its bandit analogue, obtaining regret bounds that characterize how weights interact with delayed feedback in both settings. Combining these bounds with our schedulers yields regret guarantees for capacity-constrained OCO under convex and strongly convex losses, for both first-order and bandit feedback. For first-order feedback, capacity $C = \Omega(\log T)$ suffices to recover the standard delayed OCO rates up to logarithmic factors. For bandit feedback, the standard delayed BCO rates are instead modulated by $(1 + \sigma_{\text{max}}/C)$ factors, where $\sigma_{\text{max}}$ is the maximum number of pending observations. This allows the regret bound to degrade gracefully when $C < \sigma_{\text{max}}$, while remaining sublinear.
Chat is not available.
Successful Page Load