paper-with-me

홈 › Papers

Multimodal Deep Learning for Robust RGB-D Object Recognition

2015-07-24 · Andreas Eitel, Jost Tobias Springenberg, Luciano Spinello, Martin Riedmiller, Wolfram Burgard

Robust object recognition is a crucial ingredient of many, if not all, real-world robotics applications. This paper leverages recent progress on Convolutional Neural Networks (CNNs) and proposes a novel RGB-D architecture for object recognition. Our architecture is composed of two separate CNN processing streams - one for each modality - which are consecutively combined with a late fusion network. We focus on learning with imperfect sensor data, a typical problem in real-world robotics tasks. For accurate learning, we introduce a multi-stage training methodology and two crucial ingredients for handling depth data with CNNs. The first, an effective encoding of depth information for CNNs that enables learning without the need for large depth datasets. The second, a data augmentation scheme for robust learning with depth images by corrupting them with realistic noise patterns. We present state-of-the-art results on the RGB-D object dataset and show recognition in challenging RGB-D real-world noisy settings.

📄 PDF Abstract BibTeX arXiv:1507.06821

Code (2)

athar71/Object_Detection_Enhancement_with_RBG-D_Data tf
isrugeek/robotics_recognition tf

Tasks

Data AugmentationDeep LearningMultimodal Deep LearningObjectObject Recognition

Similar Papers 제목 키워드 기반

Text-driven object affordance for guiding grasp-type recognition in multimodal robot teaching

2021-02-27 · Naoki Wake, Daichi Saito, Kazuhiro Sasabuchi, Hideki Koike 외

This study investigates how text-driven object affordance, which provides prior knowledge about grasp types for each object, affects image-based grasp-type recognition in robot teaching. The researchers created labeled d…

Mixed RealityObjectVocal Bursts Type Prediction

Hierarchical Multimodal Metric Learning for Multimodal Classification

2017-07-01 · CVPR 2017 7 · Heng Zhang, Vishal M. Patel, Rama Chellappa

Multimodal classification arises in many computer vision tasks such as object classification and image retrieval. The idea is to utilize multiple sources (modalities) measuring the same instance to improve the overall pe…

ClassificationGeneral ClassificationImage RetrievalMetric Learning+2

Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition

2026-01-30 · Sumana Biswas, Karen Young, Josephine Griffith arxiv

Multimodal sentiment analysis, which includes both image and text data, presents several challenges due to the dissimilarities in the modalities of text and image, the ambiguity of sentiment, and the complexities of cont…

Multimodal Sentiment AnalysisObject Recognition

The Labeled Multiple Canonical Correlation Analysis for Information Fusion

2021-02-28 · Lei Gao, Rui Zhang, Lin Qi, Enqing Chen 외

The objective of multimodal information fusion is to mathematically analyze information carried in different sources and create a new representation which will be more effectively utilized in pattern recognition and othe…

Emotion RecognitionFace RecognitionHandwritten Digit RecognitionObject Recognition

iCub! Do you recognize what I am doing?: multimodal human action recognition on multisensory-enabled iCub robot

2022-12-17 · Kas Kniesmeijer, Murat Kirtay

This study uses multisensory data (i.e., color and depth) to recognize human actions in the context of multimodal human-robot interaction. Here we employed the iCub robot to observe the predefined actions of the human pa…

Action RecognitionEnsemble LearningTemporal Action Localization