paper-with-me

홈 › Papers

When does dough become a bagel? Analyzing the remaining mistakes on ImageNet

2022-05-09 · Vijay Vasudevan, Benjamin Caine, Raphael Gontijo-Lopes, Sara Fridovich-Keil, Rebecca Roelofs

Image classification accuracy on the ImageNet dataset has been a barometer for progress in computer vision over the last decade. Several recent papers have questioned the degree to which the benchmark remains useful to the community, yet innovations continue to contribute gains to performance, with today's largest models achieving 90%+ top-1 accuracy. To help contextualize progress on ImageNet and provide a more meaningful evaluation for today's state-of-the-art models, we manually review and categorize every remaining mistake that a few top models make in order to provide insight into the long-tail of errors on one of the most benchmarked datasets in computer vision. We focus on the multi-label subset evaluation of ImageNet, where today's best models achieve upwards of 97% top-1 accuracy. Our analysis reveals that nearly half of the supposed mistakes are not mistakes at all, and we uncover new valid multi-labels, demonstrating that, without careful review, we are significantly underestimating the performance of these models. On the other hand, we also find that today's best models still make a significant number of mistakes (40%) that are obviously wrong to human reviewers. To calibrate future progress on ImageNet, we provide an updated multi-label evaluation set, and we curate ImageNet-Major: a 68-example "major error" slice of the obvious mistakes made by today's top models -- a slice where models should achieve near perfection, but today are far from doing so.

📄 PDF Abstract BibTeX arXiv:2205.04596

Code (1)

google-research/imagenet-mistakes 공식 구현 tf

Tasks

image-classificationImage Classification

Similar Papers 제목 키워드 기반

Robotic Dough Shaping

2022-07-31 · Jan Ondras, Di Ni, Xi Deng, Zeqi Gu 외

Robotic manipulation of deformable objects gains great attention due to its wide applications including medical surgery, home assistance, and automatic food preparation. The ability to deform soft objects remains a great…

Sand

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

2026-07-03 · Zelin Zhao, Min Shi, Bo Yuan, Haotian Xue 외 arxiv

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-World (VLAW) modeling. Meanwhile, unified vi…

multimodal generation

DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects

2024-04-18 · Dominik Bauer, Zhenjia Xu, Shuran Song

Manipulation of elastoplastic objects like dough often involves topological changes such as splitting and merging. The ability to accurately predict these topological changes that a specific action might incur is critica…

Denoising

Modelling the Doughnut of social and planetary boundaries with frugal machine learning

2025-12-01 · Stefano Vrizzi, Daniel W. O'Neill arxiv

The 'Doughnut' of social and planetary boundaries has emerged as a popular framework for assessing environmental and social sustainability. Here, we provide a proof-of-concept analysis that shows how machine learning (ML…

Reinforcement Learning

Pattern Attention Transformer with Doughnut Kernel

2022-11-30 · Wenyuan Sheng

We present in this paper a new architecture, the Pattern Attention Transformer (PAT), that is composed of the new doughnut kernel. Compared with tokens in the NLP field, Transformer in computer vision has the problem of …

image-classificationImage Classification