Recurrent Models for Situation Recognition
This work proposes Recurrent Neural Network (RNN) models to predict structured 'image situations' -- actions and noun entities fulfilling semantic roles related to the action. In contrast to prior work relying on Conditional Random Fields (CRFs), we use a specialized action prediction network followed by an RNN for noun prediction. Our system obtains state-of-the-art accuracy on the challenging recent imSitu dataset, beating CRF-based models, including ones trained with additional data. Further, we show that specialized features learned from situation prediction can be transferred to the task of image captioning to more accurately describe human-object interactions.
Code (0)
등록된 구현이 없습니다.
Tasks
Grounded Situation RecognitionHuman-Object Interaction DetectionImage CaptioningPredictionSituation RecognitionSimilar Papers 제목 키워드 기반
An Ontology Design Pattern for representing Recurrent Situations
In this paper, we present an Ontology Design Pattern for representing situations that recur at regular periods and share some invariant factors, which unify them conceptually: we refer to this set of recurring situations…
Applications of Recurrent Neural Network for Biometric Authentication & Anomaly Detection
Recurrent Neural Networks are powerful machine learning frameworks that allow for data to be saved and referenced in a temporal sequence. This opens many new possibilities in fields such as handwriting analysis and speec…
Anomaly Detectionspeech-recognitionSpeech RecognitionLearning Multi-level Dependencies for Robust Word Recognition
Robust language processing systems are becoming increasingly important given the recent awareness of dangerous situations where brittle machine learning models can be easily broken with the presence of noises. In this pa…
Position and Rotation Invariant Sign Language Recognition from 3D Kinect Data with Recurrent Neural Networks
Sign language is a gesture-based symbolic communication medium among speech and hearing impaired people. It also serves as a communication bridge between non-impaired and impaired populations. Unfortunately, in most situ…
Sign Language RecognitionEMERSK -- Explainable Multimodal Emotion Recognition with Situational Knowledge
Automatic emotion recognition has recently gained significant attention due to the growing popularity of deep learning algorithms. One of the primary challenges in emotion recognition is effectively utilizing the various…
Emotion RecognitionMultimodal Emotion Recognition