Object Recognition
9개 벤치마크 · 논문 2,197편 · 이 태스크의 논문 보기 →
Benchmarks
shape bias
CIFAR10-DVS
N-Caltech 101
ObjectNet (All classes)
DVS128 Gesture
MECCANO
N-CARS
Most implemented
Densely Connected Convolutional Networks
A Simple Framework for Contrastive Learning of Visual Representations
Going Deeper with Convolutions
Learning Transferable Visual Models From Natural Language Supervision
Microsoft COCO: Common Objects in Context
Striving for Simplicity: The All Convolutional Net
Papers
Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions
Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning…
Object RecognitionKnowledge GraphsTexture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks
Texture image classification plays a significant role in computer vision applications, including industrial inspection, medical image analysis, remote sensing, and object recognition. Handcrafted features can capture loc…
Image ClassificationObject RecognitionEgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate imp…
Binary ClassificationObject RecognitionSAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, espe…
Object RecognitionRobot ManipulationPoint CloudsCRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…
Visual Question AnsweringReferring ExpressionObject RecognitionVisual GroundingSTSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex
The primate visual system is typically divided into two streams - the ventral stream, responsible for object recognition, and the dorsal stream, responsible for encoding spatial relations and motion. Recent studies have …
Object Recognition