KINDER: Kernel-based Independence for Fair Representation Learning via Prototype-space Erasure
Abstract
Concept erasure is a prominent approach to achieving fairness in machine learning through the removal of sensitive attributes from learned representations. Prior concept erasure methods typically define debiasing indirectly through the failure of a chosen adversary, probe, or fairness penalty, making the target of erasure dependent on a particular decoder family or optimization setup. Rather than depending on a particular architecture, task type, or training paradigm, we introduce KINDER, which aims to directly remove sensitive-attribute information from a model's intermediate representations and is compatible with any deep model that contains an intermediate feature space. KINDER operates in a random Fourier feature space, where it estimates a sensitive subspace from protected group prototypes and projects representations onto its orthogonal complement, thereby forming an explicit representation cleaning mechanism. We conduct extensive experiments across diverse settings and show the effectiveness of KINDER on (i) unimodal and multimodal datasets, (ii) supervised and self-supervised settings, (iii) classification, regression, and image segmentation tasks, and (iv) diverse data modalities, including visual, textual, and tabular data.