LACE: Latent Alignment via Counterfactual Embeddings
Abstract
Improving compositional understanding in CLIP-style vision-language models remains difficult, models with strong coarse alignment often fail on attribute binding, relations, and multi-object semantics. A common remedy is text-side hard negatives (negations, perturbations, or mined confusions), but concentrating difficulty on the text branch can skew the training signal, exacerbate modality asymmetry, and distort joint-space geometry (e.g., increased hubness and reduced mutual reciprocity), even when Recall@K improves. We propose LACE (Latent Alignment via Counterfactual Embeddings), a lightweight approach that injects structured, semantics-preserving latent edits into the image and text embedding streams without pixel-space editing or external generators. LACE synthesizes counterfactual image and text embeddings by attenuating, swapping, or transplanting a small set of factors tied to target objects, attributes, or relations, producing image-side and text-side hard negatives that rebalance supervision while preserving global alignment. Across compositional and retrieval benchmarks, LACE improves robustness to binding and relational confusions and yields a healthier embedding geometry with improved reciprocity and competitive hubness/local-stability trade-offs.