Contextual Translation Embedding for Visual Relationship Detection and Scene Graph Generation
Relations amongst entities play a central role in image understanding. Due to the complexity of modeling (subject, predicate, object) relation triplets, it is crucial to develop a method that can not only recognize seen relations, but also generalize to unseen cases. Inspired by a previously proposed visual translation embedding model, or VTransE, we propose a context-augmented translation embedding model that can capture both common and rare relations. The previous VTransE model maps entities and predicates into a low-dimensional embedding vector space where the predicate is interpreted as a translation vector between the embedded features of the bounding box regions of the subject and the object. Our model additionally incorporates the contextual information captured by the bounding box of the union of the subject and the object, and learns the embeddings guided by the constraint predicate $\approx$ union (subject, object) $-$ subject $-$ object. In a comprehensive evaluation on multiple challenging benchmarks, our approach outperforms previous translation-based models and comes close to or exceeds the state of the art across a range of settings, from small-scale to large-scale datasets, from common to previously unseen relations. It also achieves promising results for the recently introduced task of scene graph generation.
Code (0)
등록된 구현이 없습니다.
Tasks
Graph GenerationObjectRelationship DetectionScene Graph GenerationTranslationVisual Relationship DetectionSimilar Papers 제목 키워드 기반
Deeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship Detection
Detecting visual relationships, i.e. <Subject, Predicate, Object> triplets, is a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We int…
Relationship DetectionScene UnderstandingTranslationVisual Relationship DetectionVisual Translation Embedding Network for Visual Relation Detection
Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language. H…
Objectobject-detectionObject DetectionRelation+4Jointly Modeling Embedding and Translation to Bridge Video and Language
Automatically describing video content with natural language is a fundamental challenge of multimedia. Recurrent Neural Networks (RNN), which models sequence dynamics, has attracted increasing attention on visual interpr…
SentenceTranslationMulti-task Learning Using a Combination of Contextualised and Static Word Embeddings for Arabic Sarcasm Detection and Sentiment Analysis
Sarcasm detection and sentiment analysis are important tasks in Natural Language Understanding. Sarcasm is a type of expression where the sentiment polarity is flipped by an interfering factor. In this study, we exploite…
Multi-Task LearningNatural Language UnderstandingSarcasm DetectionSentiment Analysis+1Detecting Concrete Visual Tokens for Multimodal Machine Translation
The challenge of visual grounding and masking in multimodal machine translation (MMT) systems has encouraged varying approaches to the detection and selection of visually-grounded text tokens for masking. We introduce ne…
Machine TranslationMultimodal Machine Translationobject-detectionObject Detection+2