paper-with-me

홈 › Papers

A comparison between humans and AI at recognizing objects in unusual poses

2024-02-06 · Netta Ollikka, Amro Abbas, Andrea Perin, Markku Kilpeläinen, Stéphane Deny

Deep learning is closing the gap with human vision on several object recognition benchmarks. Here we investigate this gap for challenging images where objects are seen in unusual poses. We find that humans excel at recognizing objects in such poses. In contrast, state-of-the-art deep networks for vision (EfficientNet, SWAG, ViT, SWIN, BEiT, ConvNext) and state-of-the-art large vision-language models (Claude 3.5, Gemini 1.5, GPT-4) are systematically brittle on unusual poses, with the exception of Gemini showing excellent robustness in that condition. As we limit image exposure time, human performance degrades to the level of deep networks, suggesting that additional mental processes (requiring additional time) are necessary to identify objects in unusual poses. An analysis of error patterns of humans vs. networks reveals that even time-limited humans are dissimilar to feed-forward deep networks. In conclusion, our comparison reveals that humans and deep networks rely on different mechanisms for recognizing objects in unusual poses. Understanding the nature of the mental processes taking place during extra viewing time may be key to reproduce the robustness of human vision in silico.

📄 PDF Abstract BibTeX arXiv:2402.03973

Code (1)

brain-aalto/unusual_poses 공식 구현

Tasks

Object Recognition

Similar Papers 제목 키워드 기반

What's Wrong With That Object? Identifying Images of Unusual Objects by Modelling the Detection Score Distribution

2016-06-01 · CVPR 2016 6 · Peng Wang, Lingqiao Liu, Chunhua Shen, Zi Huang 외

This paper studies the challenging problem of identifying unusual instances of known objects in images within an "open world" setting. That is, we aim to find objects that are members of a known class, but which are not …

Gaussian ProcessesObjectobject-detectionObject Detection

Learning Abstract Classes using Deep Learning

2016-06-17 · Sebastian Stabinger, Antonio Rodriguez-Sanchez, Justus Piater

Humans are generally good at learning abstract concepts about objects and scenes (e.g.\ spatial orientation, relative sizes, etc.). Over the last years convolutional neural networks have achieved almost human performance…

Deep Learning

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

2025-05-11 · Xilin Jiang, Junkai Wu, Vishal Choudhari, Nima Mesgarani

Audio large language models (LLMs) are considered experts at recognizing sound objects, yet their performance relative to LLMs in other sensory modalities, such as visual or audio-visual LLMs, and to humans using their e…

Transfer Learning

Quantum learning and essential cognition under the traction of meta-characteristics in an open world

2023-11-22 · Jin Wang, Changlin Song

Artificial intelligence has made significant progress in the Close World problem, being able to accurately recognize old knowledge through training and classification. However, AI faces significant challenges in the Open…

Motion Guided Attention Fusion to Recognize Interactions from Videos

2021-04-01 · ICCV 2021 10 · Tae Soo Kim, Jonathan Jones, Gregory D. Hager

We present a dual-pathway approach for recognizing fine-grained interactions from videos. We build on the success of prior dual-stream approaches, but make a distinction between the static and dynamic representations of …

Action Recognitionobject-detectionObject Detection