COHE: Auditing Non-Transitivity in Sample Difficulty Proxies for Vision Models
Abstract
Label-free proxy scores are often used as if they measured downstream discriminative difficulty. A score may rank samples similarly to a target, predict that target on held-out data, or improve a downstream decision; these are distinct evidence levels. We present Cross-Objective Hardness Evaluation (COHE), a protocol that audits each proxy--target--protocol tuple through dependence, held-out-prediction, and operational-transfer gates. COHE maps standard, inspectable statistics to bounded claim tiers; the contribution is evidence-to-claim calibration, not a new sampler, metric, or model. A DINOv2 isolation stress test illustrates the boundary: representation isolation captures a first-learning dynamics relation (rho = 0.4363, R^2 = 0.2126), but using it for direct 10% subset selection fails against Random (Delta = -15.41, 95% CI [-15.87, -14.95]). COHE therefore preserves this target-specific relation without rebranding it as actionability. Across CIFAR-10/100, ImageNet-10, and ImageNet-1K validation checks, eight primary generative proxy families do not support CE/margin surrogate claims: they show weak alignment, near-zero held-out CE prediction, and unreliable transfer. Controls show stable discriminative target rankings, expected shuffled-proxy null behavior, and a task-aligned ImageNet target-hard retrieval/triage tuple that passes all gates. The anonymized supplementary package provides lightweight gate-wise audit code, toy scalar CSV examples, cached aggregate summaries, selected derived scalar arrays for tabular gate verification, supported paper-style verification scripts, and provenance notes without redistributing raw datasets, images, weights, checkpoints, or full scoring/retraining pipelines.