Joint Object and State Recognition using Language Knowledge
The state of an object is an important piece of knowledge in robotics applications. States and objects are intertwined together, meaning that object information can help recognize the state of an image and vice versa. This paper addresses the state identification problem in cooking related images and uses state and object predictions together to improve the classification accuracy of objects and their states from a single image. The pipeline presented in this paper includes a CNN with a double classification layer and the Concept-Net language knowledge graph on top. The language knowledge creates a semantic likelihood between objects and states. The resulting object and state confidences from the deep architecture are used together with object and state relatedness estimates from a language knowledge graph to produce marginal probabilities for objects and states. The marginal probabilities and confidences of objects (or states) are fused together to improve the final object (or state) classification results. Experiments on a dataset of cooking objects show that using a language knowledge graph on top of a deep neural network effectively enhances object and state classification.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationGeneral ClassificationObjectSimilar Papers 제목 키워드 기반
CKERC : Joint Large Language Models with Commonsense Knowledge for Emotion Recognition in Conversation
Emotion recognition in conversation (ERC) is a task which predicts the emotion of an utterance in the context of a conversation. It tightly depends on dialogue context, speaker identity information, multiparty dialogue s…
Emotion RecognitionEmotion Recognition in ConversationLanguage ModelingLanguage Modelling+1Verb Physics: Relative Physical Knowledge of Actions and Objects
Learning commonsense knowledge from natural language text is nontrivial due to reporting bias: people rarely state the obvious, e.g., "My house is bigger than me." However, while rarely stated explicitly, this trivial ev…
H2O: Two Hands Manipulating Objects for First Person Interaction Recognition
We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3…
Action Recognitionhand-object poseObjectPose Estimation+1Independent Sign Language Recognition with 3D Body, Hands, and Face Reconstruction
Independent Sign Language Recognition is a complex visual recognition problem that combines several challenging tasks of Computer Vision due to the necessity to exploit and fuse information from hand gestures, body featu…
3D Action Recognition3D ReconstructionAction RecognitionFace Reconstruction+2Bowtie Networks: Generative Modeling for Joint Few-Shot Recognition and Novel-View Synthesis
We propose a novel task of joint few-shot recognition and novel-view synthesis: given only one or few images of a novel object from arbitrary views with only category annotation, we aim to simultaneously learn an object …
Data AugmentationMulti-Task LearningNovel View SynthesisObject