Ask the Model: Counterfactual Cue Reliance Predicts Worst-Group Failure and Targets Repair
Abstract
Fine-tuned models can rely on shortcut cues that correlate with labels during training but mislead on some inputs. We ask whether a finished model's behavior reveals how much it relies on such a cue. Given a suspected cue, we measure counterfactual cue reliance, Δcue, by editing the cue and measuring the change in the model's decision score. This requires no group labels, training data, or knowledge of the training recipe. We test 486 fine-tuned language models on CivilComments and Bias-in-Bios while varying the training cue--label correlation. Across 180 models in 10 main task--model settings, higher Δcue predicts lower worst-group accuracy. In a separate high-seed analysis, Δcue predicts at held-out correlation levels, where linear extrapolation from the assigned correlation performs worse than predicting the mean. Capability and placebo-edit controls argue against model quality and generic edit sensitivity as explanations. Δcue also provides signal for prioritizing concept-erasure repair. Exact control shows that most of the predictive signal comes from differences across correlation levels. Within a fixed recipe, the relationship across random seeds is small and not robust.