paper-with-me

홈 › Papers

Humans and deep networks largely agree on which kinds of variation make object recognition harder

2016-04-21 · Saeed Reza Kheradpisheh, Masoud Ghodrati, Mohammad Ganjtabesh, Timothée Masquelier

View-invariant object recognition is a challenging problem, which has attracted much attention among the psychology, neuroscience, and computer vision communities. Humans are notoriously good at it, even if some variations are presumably more difficult to handle than others (e.g. 3D rotations). Humans are thought to solve the problem through hierarchical processing along the ventral stream, which progressively extracts more and more invariant visual features. This feed-forward architecture has inspired a new generation of bio-inspired computer vision systems called deep convolutional neural networks (DCNN), which are currently the best algorithms for object recognition in natural images. Here, for the first time, we systematically compared human feed-forward vision and DCNNs at view-invariant object recognition using the same images and controlling for both the kinds of transformation as well as their magnitude. We used four object categories and images were rendered from 3D computer models. In total, 89 human subjects participated in 10 experiments in which they had to discriminate between two or four categories after rapid presentation with backward masking. We also tested two recent DCNNs on the same tasks. We found that humans and DCNNs largely agreed on the relative difficulties of each kind of variation: rotation in depth is by far the hardest transformation to handle, followed by scale, then rotation in plane, and finally position. This suggests that humans recognize objects mainly through 2D template matching, rather than by constructing 3D object models, and that DCNNs are not too unreasonable models of human feed-forward vision. Also, our results show that the variation levels in rotation in depth and scale strongly modulate both humans' and DCNNs' recognition performances. We thus argue that these variations should be controlled in the image datasets used in vision research.

📄 PDF Abstract BibTeX arXiv:1604.06486

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject RecognitionTemplate Matching

Similar Papers 제목 키워드 기반

Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains

2026-04-19 · Finn Schmidt, Jan Philip Wahle, Terry Ruas, Bela Gipp arxiv

Automatic evaluation metrics are central to the development of machine translation systems, yet their robustness under domain shift remains unclear. Most metrics are developed on the Workshop on Machine Translation (WMT)…

Machine Translation

Mechanisms for Handling Nested Dependencies in Neural-Network Language Models and Humans

2020-06-19 · Yair Lakretz, Dieuwke Hupkes, Alessandra Vergallito, Marco Marelli 외

Recursive processing in sentence comprehension is considered a hallmark of human linguistic abilities. However, its underlying neural mechanisms remain largely unknown. We studied whether a modern artificial neural netwo…

Sentence

The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions

2026-08-06 · Hadi Hosseini, Samarth Khanna, Leona Pierce arxiv

As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, pa…

Annotator Response Distributions as a Sampling Frame

2022-06-01 · NLPerspectives (LREC) 2022 6 · Christopher Homan, Tharindu Cyril Weerasooriya, Lora Aroyo, Chris Welty

Annotator disagreement is often dismissed as noise or the result of poor annotation process quality. Others have argued that it can be meaningful. But lacking a rigorous statistical foundation, the analysis of disagreeme…

Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting

2023-09-13 · Tilman Beck, Hendrik Schuff, Anne Lauscher, Iryna Gurevych

Annotators' sociodemographic backgrounds (i.e., the individual compositions of their gender, age, educational background, etc.) have a strong impact on their decisions when working on subjective NLP tasks, such as toxic …

Hate Speech DetectionSensitivityZero-Shot Learning