Visual Entailment Task for Visually-Grounded Language Learning
We introduce a new inference task - Visual Entailment (VE) - which differs from traditional Textual Entailment (TE) tasks whereby a premise is defined by an image, rather than a natural language sentence as in TE tasks. A novel dataset SNLI-VE (publicly available at https://github.com/necla-ml/SNLI-VE) is proposed for VE tasks based on the Stanford Natural Language Inference corpus and Flickr30k. We introduce a differentiable architecture called the Explainable Visual Entailment model (EVE) to tackle the VE problem. EVE and several other state-of-the-art visual question answering (VQA) based models are evaluated on the SNLI-VE dataset, facilitating grounded language understanding and providing insights on how modern VQA based models perform.
Code (1)
Tasks
Grounded language learningNatural Language InferenceQuestion AnsweringSentenceVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Grounded Textual Entailment
Capturing semantic relations between sentences, such as entailment, is a long-standing challenge for computational semantics. Logic-based models analyse entailment in terms of possible worlds (interpretations, or situati…
Natural Language InferenceVisually grounded generation of entailments from premises
Natural Language Inference (NLI) is the task of determining the semantic relationship between a premise and a hypothesis. In this paper, we focus on the generation of hypotheses from premises in a multimodal setting, to …
Natural Language InferenceSentenceMultimodal Logical Inference System for Visual-Textual Entailment
A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations. In this paper, we use logic-based representations as…
Automated Theorem ProvingNatural Language InferenceSemantic ParsingSentenceVisually-Verifiable Textual Entailment: A Challenge Task for Combining Language and Vision
Using Grounded Word Representations to Study Theories of Lexical Concepts
The fields of cognitive science and philosophy have proposed many different theories for how humans represent {``}concepts{''}. Multiple such theories are compatible with state-of-the-art NLP methods, and could in princi…
Philosophy