Doomed to Re-Annotate, Forever: The ImageNet Story
Illia Volkov ⋅ Nikita Kisel ⋅ Tetiana Mishkina ⋅ Klara Janouskova ⋅ Jiri Matas
Abstract
Top-1 accuracy on ImageNet-1k remains the most universally reported metric in visual recognition, despite its well-documented label quality issues. In the paper, we describe the effort to obtain accurate and complete ImageNet-1k validation set annotations. The result -- ReImageNet -- is a from-scratch reannotation, which includes multilabel correction, object localisation, revised class definitions, and semantic attributes---text-recognition, rendition, reflection, crowd, and dominant. The reannotation reveals that $\approx13$\% of the original ImageNet-1k labels are incorrect, $32.7$\% of images are multilabel and $4.7$\% contain no object from an ImageNet-1k class. With the new labels, top-1 accuracy increases by $0.7$--$1.4$\% for supervised models and by $6$--$8$\% for MLLMs. We show that annotation errors in ImageNet-1k propagate into its derivative test sets, indicating that the problem is structural rather than specific to any single benchmark. We observed that LLMs outperform unaided crowd annotators, and that human and LLM collaboration with appropriate tooling represents the current quality ceiling for annotation at this scale. All [annotations](https://huggingface.co/datasets/c1rcuslegend/ReImageNet), [class definitions](https://gonikisgo.github.io/imagenet-annotation-preview/), [guidelines](https://gonikisgo.github.io/imagenet-annotation-preview/), and [analysis code](https://github.com/klarajanouskova/ImageNet) have been publicly released.
Chat is not available.
Successful Page Load