Papers Object Categorization
“Object Categorization” 태그가 달린 논문 80편 · 필터 해제
Vision CNNs trained to estimate spatial latents learned similar ventral-stream-aligned representations
Studies of the functional role of the primate ventral visual stream have traditionally focused on object categorization, often ignoring -- despite much prior evidence -- its role in estimating "spatial" latents such as o…
Object CategorizationDivide and Conquer: Improving Multi-Camera 3D Perception with 2D Semantic-Depth Priors and Input-Dependent Queries
3D perception tasks, such as 3D object detection and Bird's-Eye-View (BEV) segmentation using multi-camera images, have drawn significant attention recently. Despite the fact that accurately estimating both semantic and …
3D Object DetectionBEV SegmentationObjectObject Categorization+3Comparing Apples to Oranges: LLM-powered Multimodal Intention Prediction in an Object Categorization Task
Human intention-based systems enable robots to perceive and interpret user actions to interact with humans and adapt to their behavior proactively. Therefore, intention prediction is pivotal in creating a natural interac…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Object Categorizationspeech-recognition+2Adversarial alignment: Breaking the trade-off between the strength of an attack and its relevance to human perception
Deep neural networks (DNNs) are known to have a fundamental sensitivity to adversarial attacks, perturbations of the input that are imperceptible to humans yet powerful enough to change the visual decision of a model. Ad…
Adversarial AttackAdversarial RobustnessDiagnosticObject Categorization+2Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers
Disaggregated performance metrics across demographic groups are a hallmark of fairness assessments in computer vision. These metrics successfully incentivized performance improvements on person-centric tasks such as face…
Fairnessimage-classificationImage ClassificationMulti-Label Image Classification+1Vocabulary-informed Zero-shot and Open-set Learning
Despite significant progress in object categorization, in recent years, a number of important challenges remain; mainly, the ability to learn from limited labeled data and to recognize object classes within large, potent…
Object CategorizationOpen Set LearningZero-Shot LearningRoboflow 100: A Rich, Multi-Domain Object Detection Benchmark
The evaluation of object detection models is usually performed by optimizing a single metric, e.g. mAP, on a fixed set of datasets, e.g. Microsoft COCO and Pascal VOC. Due to image retrieval and annotation costs, these d…
2D Object DetectionImage RetrievalMedical Object DetectionMulti-Object Tracking+14Enhancing Fine-Grained 3D Object Recognition using Hybrid Multi-Modal Vision Transformer-CNN Models
Robots operating in human-centered environments, such as retail stores, restaurants, and households, are often required to distinguish between similar objects in different contexts with a high degree of accuracy. However…
3D Object RecognitionFine-Grained Image ClassificationObject CategorizationObject RecognitionUnified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and image generation, vision-and-language t…
Depth EstimationImage GenerationKeypoint EstimationObject Categorization+10GRIT: General Robust Image Task Benchmark
Computer vision models excel at making predictions when the test distribution closely resembles the training distribution. Such models have yet to match the ability of biological vision to learn from multiple sources and…
Instance SegmentationKeypoint DetectionObject CategorizationObject Localization+6OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
In this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization. We propose OFA, a Task-Agnostic and Modality-Agnostic framework that supports …
Image Captioningimage-classificationImage ClassificationImage Generation+14Webly Supervised Concept Expansion for General Purpose Vision Models
General Purpose Vision (GPV) systems are models that are designed to solve a wide array of visual tasks without requiring architectural changes. Today, GPVs primarily learn both skills and concepts from large fully super…
Human-Object Interaction DetectionImage RetrievalObject CategorizationObject Localization+2Category-orthogonal object features guide information processing in recurrent neural networks trained for object categorization
Recurrent neural networks (RNNs) have been shown to perform better than feedforward architectures in visual object categorization tasks, especially in challenging conditions such as cluttered images. However, little is k…
DiagnosticObjectObject CategorizationLearning Transferable Visual Models From Natural Language Supervision
State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories. This restricted form of supervision limits their generality and usability since additional labeled data is n…
Action RecognitionBenchmarkingFew-Shot Image Classificationgeo-localization+21Open-Ended Fine-Grained 3D Object Categorization by Combining Shape and Texture Features in Multiple Colorspaces
As a consequence of an ever-increasing number of service robots, there is a growing demand for highly accurate real-time 3D object recognition. Considering the expansion of robot applications in more complex and dynamic …
3D Object RecognitionDescriptiveObjectObject Categorization+2Local-HDP: Interactive Open-Ended 3D Object Categorization in Real-Time Robotic Scenarios
We introduce a non-parametric hierarchical Bayesian approach for open-ended 3D object categorization, named the Local Hierarchical Dirichlet Process (Local-HDP). This method allows an agent to learn independent topics fo…
Object CategorizationVariational InferenceIAUnet: Global Context-Aware Feature Learning for Person Re-Identification
Person re-identification (reID) by CNNs based networks has achieved favorable performance in recent years. However, most of existing CNNs based methods do not take full advantage of spatial-temporal context modeling. In …
Object CategorizationPerson Re-IdentificationLearning Physical Graph Representations from Visual Scenes
Convolutional Neural Networks (CNNs) have proved exceptional at learning representations for visual object categorization. However, CNNs do not explicitly encode objects, parts, and their physical properties, which has l…
ObjectObject CategorizationScene SegmentationUnsupervised Domain Adaptation through Inter-modal Rotation for RGB-D Object Recognition
Unsupervised Domain Adaptation (DA) exploits the supervision of a label-rich source dataset to make predictions on an unlabeled target dataset by aligning the two data distributions. In robotics, DA is used to take advan…
Domain AdaptationObject CategorizationObject RecognitionUnsupervised Domain AdaptationBrain-Like Object Recognition with High-Performing Shallow Recurrent ANNs
Deep convolutional artificial neural networks (ANNs) are the leading class of candidate models of the mechanisms of visual processing in the primate ventral stream. While initially inspired by brain anatomy, over the pas…
AnatomyBIG-bench Machine LearningObject CategorizationObject Recognition+1