Adversarial Group Fairness in Contextual Bandits: When Robust is Not Fair
Ray Telikani ⋅ Jaber valizadeh ⋅ Amir H Gandomi ⋅ Bao Q Vo ⋅ Ming Ding
Abstract
Neural contextual bandits support high-stakes decision systems---from clinical trial allocation to content recommendation. This paper investigates an adversarial threat, group-targeted suppression attacks (GTSA), in which an adversary selectively degrades performance for a demographic subgroup while preserving global performance metrics. We first show that under standard aggregate monitoring, GTSA with budget $B = \Omega(\sqrt{T})$ can induce $\Omega(T)$ group regret while maintaining only $o(1)$ deviation in observable global statistics, establishing that any defense that ignores group structure is fundamentally vulnerable. To address this vulnerability, we propose AnchorFair, a robust mitigation framework comprising three components: (i) a trust-weighted policy anchor with provable bounded drift, (ii) representation-stable demographic discovery via online clustering in a slowly evolving embedding space, and (iii) a demographic-aware exploration mechanism that adaptively amplifies learning for underrepresented or low-trust groups. We prove that \textsc{AnchorFair} achieves sublinear group regret $\tilde{O}(\sqrt{T})$ for all groups under GTSA.
Chat is not available.
Successful Page Load