paper-with-me

Papers Object Categorization

“Object Categorization” 태그가 달린 논문 80편 · 필터 해제

Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representations

2024-12-12 · Yudi Xie, Weichen Huang, Esther Alter, Jeremy Schwartz 외

Studies of the functional role of the primate ventral visual stream have traditionally focused on object categorization, often ignoring -- despite much prior evidence -- its role in estimating "spatial" latents such as o…

Object Categorization

Divide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries

2024-08-13 · Qi Song, Qingyong Hu, Chi Zhang, Yongquan Chen 외

3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that accurately estimating both semantic and …

3D Object DetectionBEV SegmentationObjectObject Categorization+3

Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task

2024-04-12 · Hassan Ali, Philipp Allgeuer, Stefan Wermter

Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interac…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object Categorizationspeech-recognition+2

Adversarial alignment: Breaking the trade-off between the strength of an attack and its relevance to human perception

2023-06-05 · Drew Linsley, Pinyuan Feng, Thibaut Boissin, Alekh Karkada Ashok 외

Deep neural networks (DNNs) are known to have a fundamental sensitivity to adversarial attacks, perturbations of the input that are imperceptible to humans yet powerful enough to change the visual decision of a model. Ad…

Adversarial AttackAdversarial RobustnessDiagnosticObject Categorization+2

Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers

2023-02-16 · Melissa Hall, Bobbie Chern, Laura Gustafson, Denisse Ventura 외

Disaggregated performance metrics across demographic groups are a hallmark of fairness assessments in computer vision. These metrics successfully incentivized performance improvements on person-centric tasks such as face…

Fairnessimage-classificationImage ClassificationMulti-Label Image Classification+1

Vocabulary-informed Zero-shot and Open-set Learning

2023-01-03 · Yanwei Fu, Xiaomei Wang, Hanze Dong, Yu-Gang Jiang 외

Despite significant progress in object categorization, in recent years, a number of important challenges remain; mainly, the ability to learn from limited labeled data and to recognize object classes within large, potent…

Object CategorizationOpen Set LearningZero-Shot Learning

Roboflow 100: A Rich, Multi-Domain Object Detection Benchmark

2022-11-24 · Floriana Ciaglia, Francesco Saverio Zuppichini, Paul Guerrie, Mark McQuade 외

The evaluation of object detection models is usually performed by optimizing a single metric, e.g. mAP, on a fixed set of datasets, e.g. Microsoft COCO and Pascal VOC. Due to image retrieval and annotation costs, these d…

2D Object DetectionImage RetrievalMedical Object DetectionMulti-Object Tracking+14

Enhancing Fine-Grained 3D Object Recognition using Hybrid Multi-Modal Vision Transformer-CNN Models

2022-10-03 · Songsong Xiong, Georgios Tziafas, Hamidreza Kasaei

Robots operating in human-centered environments, such as retail stores, restaurants, and households, are often required to distinguish between similar objects in different contexts with a high degree of accuracy. However…

3D Object RecognitionFine-Grained Image ClassificationObject CategorizationObject Recognition

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

2022-06-17 · Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi 외

We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and image generation, vision-and-language t…

Depth EstimationImage GenerationKeypoint EstimationObject Categorization+10

GRIT: General Robust Image Task Benchmark

2022-04-28 · Tanmay Gupta, Ryan Marten, Aniruddha Kembhavi, Derek Hoiem

Computer vision models excel at making predictions when the test distribution closely resembles the training distribution. Such models have yet to match the ability of biological vision to learn from multiple sources and…

Instance SegmentationKeypoint DetectionObject CategorizationObject Localization+6

OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework

2022-02-07 · Peng Wang, An Yang, Rui Men, Junyang Lin 외

In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports …

Image Captioningimage-classificationImage ClassificationImage Generation+14

Webly Supervised Concept Expansion for General Purpose Vision Models

2022-02-04 · Amita Kamath, Christopher Clark, Tanmay Gupta, Eric Kolve 외

General Purpose Vision (GPV) systems are models that are designed to solve a wide array of visual tasks without requiring architectural changes. Today, GPVs primarily learn both skills and concepts from large fully super…

Human-Object Interaction DetectionImage RetrievalObject CategorizationObject Localization+2

Category-orthogonal object features guide information processing in recurrent neural networks trained for object categorization

2021-11-15 · NeurIPS Workshop SVRHM 2021 12 · Sushrut Thorat, Giacomo Aldegheri, Tim C. Kietzmann

Recurrent neural networks (RNNs) have been shown to perform better than feedforward architectures in visual object categorization tasks, especially in challenging conditions such as cluttered images. However, little is k…

DiagnosticObjectObject Categorization

Learning Transferable Visual Models From Natural Language Supervision

2021-02-26 · Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 외

State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…

Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21

Open-Ended Fine-Grained 3D Object Categorization by Combining Shape and Texture Features in Multiple Colorspaces

2020-09-19 · Nils Keunecke, S. Hamidreza Kasaei

As a consequence of an ever-increasing number of service robots, there is a growing demand for highly accurate real-time 3D object recognition. Considering the expansion of robot applications in more complex and dynamic …

3D Object RecognitionDescriptiveObjectObject Categorization+2

Local-HDP: Interactive Open-Ended 3D Object Categorization in Real-Time Robotic Scenarios

2020-09-02 · H. Ayoobi, H. Kasaei, M. Cao, R. Verbrugge 외

We introduce a non-parametric hierarchical Bayesian approach for open-ended 3D object categorization, named the Local Hierarchical Dirichlet Process (Local-HDP). This method allows an agent to learn independent topics fo…

Object CategorizationVariational Inference

IAUnet: Global Context-Aware Feature Learning for Person Re-Identification

2020-09-02 · Ruibing Hou, Bingpeng Ma, Hong Chang, Xinqian Gu 외

Person re-identification (reID) by CNNs based networks has achieved favorable performance in recent years. However, most of existing CNNs based methods do not take full advantage of spatial-temporal context modeling. In …

Object CategorizationPerson Re-Identification

Learning Physical Graph Representations from Visual Scenes

2020-06-22 · NeurIPS 2020 12 · Daniel M. Bear, Chaofei Fan, Damian Mrowca, Yunzhu Li 외

Convolutional Neural Networks (CNNs) have proved exceptional at learning representations for visual object categorization. However, CNNs do not explicitly encode objects, parts, and their physical properties, which has l…

ObjectObject CategorizationScene Segmentation

Unsupervised Domain Adaptation through Inter-modal Rotation for RGB-D Object Recognition

2020-04-21 · Mohammad Reza Loghmani, Luca Robbiano, Mirco Planamente, Kiru Park 외

Unsupervised Domain Adaptation (DA) exploits the supervision of a label-rich source dataset to make predictions on an unlabeled target dataset by aligning the two data distributions. In robotics, DA is used to take advan…

Domain AdaptationObject CategorizationObject RecognitionUnsupervised Domain Adaptation

Brain-Like Object Recognition with High-Performing Shallow Recurrent ANNs

2019-09-13 · NeurIPS 2019 12 · Jonas Kubilius, Martin Schrimpf, Kohitij Kar, Ha Hong 외

Deep convolutional artificial neural networks (ANNs) are the leading class of candidate models of the mechanisms of visual processing in the primate ventral stream. While initially inspired by brain anatomy, over the pas…

AnatomyBIG-bench Machine LearningObject CategorizationObject Recognition+1
1–20 / 80 다음 →