paper-with-me

Object Recognition

9개 벤치마크 · 논문 2,197편 · 이 태스크의 논문 보기 →

Benchmarks

shape bias

결과 18개

CIFAR10-DVS

결과 2개

N-Caltech 101

결과 2개

DVS128 Gesture

결과 1개

MECCANO

결과 1개

N-CARS

결과 1개

Most implemented

Densely Connected Convolutional Networks

2016-08-25 · 구현 146개

Going Deeper with Convolutions

2014-09-17 · 구현 83개

Microsoft COCO: Common Objects in Context

2014-05-01 · 구현 38개

Papers

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

2026-09-04 · Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta 외 arxiv

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning…

Object RecognitionKnowledge Graphs

Texture Image Classification Using DWT AlexNet Feature Fusion and Deep Neural Networks

2026-08-28 · Arun D. Kulkarni arxiv

Texture image classification plays a significant role in computer vision applications, including industrial inspection, medical image analysis, remote sensing, and object recognition. Handcrafted features can capture loc…

Image ClassificationObject Recognition

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

2026-07-29 · Yuyun Chen, Tianao Li, TianQuan Feng, Cen Chen 외 arxiv

Reliable visual safety understanding in real-world scenarios demands more than just object recognition; it requires causal reasoning under epistemic uncertainty. While Large Vision-Language Models (LVLMs) demonstrate imp…

Binary ClassificationObject Recognition

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

2026-07-28 · Zonghe Liu, Shanyuan Jie, Xiaoquan Sun, Chen Cao 외 arxiv

Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, espe…

Object RecognitionRobot ManipulationPoint Clouds

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

2026-07-23 · Hanseok Oh, Parishad BehnamGhader, Benno Krojer, Hyunji Lee 외 arxiv

Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond a provided image to answer questions. K…

Visual Question AnsweringReferring ExpressionObject RecognitionVisual Grounding

STSBench: A Large-Scale Dataset for Modeling Neuronal Activity in the Dorsal Stream of Primate Visual Cortex

2026-07-17 · Ethan B. Trepka, Ruobing Xia, Shude Zhu, Sharif Saleki 외 arxiv

The primate visual system is typically divided into two streams - the ventral stream, responsible for object recognition, and the dorsal stream, responsible for encoding spatial relations and motion. Recent studies have …

Object Recognition

전체 2,197편 보기 →