Sparse Feature Compositionality in Vision-Language Models
Abstract
Sparse features in vision-language models are increasingly used as interpretable units for intervention, yet it remains unclear whether individually meaningful features remain reliable when composed. We study this question in a controlled visual reasoning setting. We localize an anchor intervention site with linear probes, train a sparse autoencoder at that site, and construct task-selective sparse feature sets. We then compare single-set and pairwise union interventions under accuracy and intervention-drift metrics. A full pairwise composition map reveals strongly heterogeneous behavior: some unions remain stable, while others substantially reduce task accuracy despite mild single-set effects. Mechanistic analysis shows that these failures are better predicted by antagonistic update geometry than by co-activation alone. The pattern persists under random, permutation, and norm-matched controls, and qualitatively transfers across datasets and model families. These results expose a reliability gap in sparse VLM intervention: feature selectivity alone does not ensure compositional control.