Correct Labels Are Not Enough: Reliable Metadata Can Silently Control Model Behavior
Abstract
Task-success evaluation indicates whether a model produced the right answer, but not which feature controlled that answer. We study this gap where supervision is perfectly correct: during supervised fine-tuning (SFT) on multiple-choice questions, the correct option always carries an inline annotation "(Reward: 5)" and incorrect options carry "(Reward: 0)". The cue is honest on every training example; the target label is never corrupted. Across Qwen3-8B, Gemma-2-9B, and Llama-3.1-8B, the honest cue becomes the feature that selects the model's answer. Moving it onto a wrong option at inference flips the answer on 97–100% of items, while cue-free accuracy stays at baseline in the CSQA setting: clean-holdout evaluation reports a healthy model, and only a counterfactual probe reveals which feature is in control. What installs the selector is the cue's reliability, not its meaning: nonsense and neutral label words install the same behavior, and sweeping the training cue's reliability reveals a threshold-like, rank-insensitive transition across four base families. The effect extends beyond SFT to direct preference optimization in our tested setting, weakening as KL regularization strengthens. Clean task success therefore certifies that a model can answer, not which learned rule produced the answer. We give a field-level counterfactual audit — move, remove, and reverse each known or suspected metadata field — for telling the two apart.