Cosine is Human: The Reproducibility Ceiling of Perceptual Similarity
Abstract
NIGHTS is a recent benchmark for human perceptual similarity over objects, scenes, and composition. DreamSim, the supervised metric introduced with NIGHTS, and later supervised metrics trained on the same data report large gains over self-supervised baselines. The standard NIGHTS evaluation uses the unanimous-agreement subset of a released 100,000-triplet corpus, leaving weak-majority and split-vote triplets outside the headline benchmark. We evaluate eleven perceptual similarity methods on the full released corpus, analyzing unanimous, weak-majority, and split-vote triplets separately. On the 64,420 non-tie triplets, cosine distance on a frozen self-supervised backbone reaches 74.1% accuracy without training, LoRA fine-tuning reaches 77–79%, and the DreamSim ensemble reaches 80.7%. A split-half reproducibility estimate from the triplet vote counts shows that majority labels reproduce only 70.1% of the time overall, and only 33–50% on weak-majority triplets. The extra accuracy comes from matching the collected votes, not from showing that the metric would better match a new set of human judgments. The evaluation target is also unstable across task formulations: NIGHTS’ 2AFC similarity labels and JND same/different labels disagree on 51–64% of shared stimuli. Fine-tuning changes the feature basis as well, increasing foreground-only fragility and making DreamSim the most fragile method under rotation, color, and blur. NIGHTS supports a narrower conclusion: cosine distance on self-supervised features captures the stable signal currently measurable in these labels. Additional claims of human alignment require a redesigned benchmark with more votes per triplet, axis-specific similarity questions, per-subject statistics, and ablation tests.