Barycentric Guidance: Turning Foundation Image Editors into Continuous Affective Controllers
Abstract
Emotions are expressed on a continuum and are naturally described in the valence-arousal (VA) space. Existing VA-based image editors are typically training-based and domain-specific, limiting them to VA regions seen during training and making them brittle across datasets and domains. Meanwhile, despite producing high-quality edits, state-of-the-art foundation image editors are not continuous affective controllers: non-linear emotion representations in text encoders makes prompt interpolation poorly calibrated, so accurate and continuous VA control remains unreliable. To fill this gap, we introduce Barycentric Guidance, a training-free framework that turns pretrained foundation image editors into continuous affective controllers over VA, requiring no finetuning, extra data, or model changes. Our method maps any target VA point to simplex-based guidance weights, enabling VA controllability in three aspects: (i) accurate targeting of VA states, (ii) broad coverage across the VA plane, and (iii) continuous control. Across three face datasets, we improve VA target accuracy by 25–29\%, expand VA-plane coverage by 28–50\%, increase expression diversity by 29–49\%, over the strongest baseline. Unlike prior domain-specific methods, our training-free approach exploits the foundation model’s inherent versatility, enabling generalization beyond human faces to animals, artworks, and complex scenes within a single framework. Code will be made publicly available.