Multimodal Deep Learning for Robust RGB-D Object Recognition
Robust object recognition is a crucial ingredient of many, if not all, real-world robotics applications. This paper leverages recent progress on Convolutional Neural Networks (CNNs) and proposes a novel RGB-D architecture for object recognition. Our architecture is composed of two separate CNN processing streams - one for each modality - which are consecutively combined with a late fusion network. We focus on learning with imperfect sensor data, a typical problem in real-world robotics tasks. For accurate learning, we introduce a multi-stage training methodology and two crucial ingredients for handling depth data with CNNs. The first, an effective encoding of depth information for CNNs that enables learning without the need for large depth datasets. The second, a data augmentation scheme for robust learning with depth images by corrupting them with realistic noise patterns. We present state-of-the-art results on the RGB-D object dataset and show recognition in challenging RGB-D real-world noisy settings.
Code (2)
Tasks
Data AugmentationDeep LearningMultimodal Deep LearningObjectObject RecognitionSimilar Papers 제목 키워드 기반
Text-driven object affordance for guiding grasp-type recognition in multimodal robot teaching
This study investigates how text-driven object affordance, which provides prior knowledge about grasp types for each object, affects image-based grasp-type recognition in robot teaching. The researchers created labeled d…
Mixed RealityObjectVocal Bursts Type PredictionHierarchical Multimodal Metric Learning for Multimodal Classification
Multimodal classification arises in many computer vision tasks such as object classification and image retrieval. The idea is to utilize multiple sources (modalities) measuring the same instance to improve the overall pe…
ClassificationGeneral ClassificationImage RetrievalMetric Learning+2Leveraging Textual-Cues for Enhancing Multimodal Sentiment Analysis by Object Recognition
Multimodal sentiment analysis, which includes both image and text data, presents several challenges due to the dissimilarities in the modalities of text and image, the ambiguity of sentiment, and the complexities of cont…
Multimodal Sentiment AnalysisObject RecognitionThe Labeled Multiple Canonical Correlation Analysis for Information Fusion
The objective of multimodal information fusion is to mathematically analyze information carried in different sources and create a new representation which will be more effectively utilized in pattern recognition and othe…
Emotion RecognitionFace RecognitionHandwritten Digit RecognitionObject RecognitioniCub! Do you recognize what I am doing?: multimodal human action recognition on multisensory-enabled iCub robot
This study uses multisensory data (i.e., color and depth) to recognize human actions in the context of multimodal human-robot interaction. Here we employed the iCub robot to observe the predefined actions of the human pa…
Action RecognitionEnsemble LearningTemporal Action Localization