paper-with-me

Papers

Visual Entailment Task for Visually-Grounded Language Learning

2018-11-26 · Ning Xie, Farley Lai, Derek Doran, Asim Kadav

We introduce a new inference task - Visual Entailment (VE) - which differs from traditional Textual Entailment (TE) tasks whereby a premise is defined by an image, rather than a natural language sentence as in TE tasks. A novel dataset SNLI-VE (publicly available at https://github.com/necla-ml/SNLI-VE) is proposed for VE tasks based on the Stanford Natural Language Inference corpus and Flickr30k. We introduce a differentiable architecture called the Explainable Visual Entailment model (EVE) to tackle the VE problem. EVE and several other state-of-the-art visual question answering (VQA) based models are evaluated on the SNLI-VE dataset, facilitating grounded language understanding and providing insights on how modern VQA based models perform.

📄 PDF Abstract BibTeX arXiv:1811.10582

Code (1)

necla-ml/SNLI-VE 공식 구현

Tasks

Grounded language learningNatural Language InferenceQuestion AnsweringSentenceVisual EntailmentVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Grounded Textual Entailment

2018-06-14 · COLING 2018 8 · Hoa Trong Vu, Claudio Greco, Aliia Erofeeva, Somayeh Jafaritazehjan 외

Capturing semantic relations between sentences, such as entailment, is a long-standing challenge for computational semantics. Logic-based models analyse entailment in terms of possible worlds (interpretations, or situati…

Natural Language Inference

Visually grounded generation of entailments from premises

2019-10-01 · WS 2019 10 · Somayeh Jafaritazehjani, Albert Gatt, Marc Tanti

Natural Language Inference (NLI) is the task of determining the semantic relationship between a premise and a hypothesis. In this paper, we focus on the generation of hypotheses from premises in a multimodal setting, to …

Natural Language InferenceSentence

Multimodal Logical Inference System for Visual-Textual Entailment

2019-06-10 · ACL 2019 7 · Riko Suzuki, Hitomi Yanaka, Masashi Yoshikawa, Koji Mineshima 외

A large amount of research about multimodal inference across text and vision has been recently developed to obtain visually grounded word and sentence representations. In this paper, we use logic-based representations as…

Automated Theorem ProvingNatural Language InferenceSemantic ParsingSentence

Visually-Verifiable Textual Entailment: A Challenge Task for Combining Language and Vision

2015-09-01 · WS 2015 9 · Jayant Krishnamurthy
Coreference ResolutionImage RetrievalNatural Language InferencePrepositional Phrase Attachment

Using Grounded Word Representations to Study Theories of Lexical Concepts

2019-06-01 · WS 2019 6 · Dylan Ebert, Ellie Pavlick

The fields of cognitive science and philosophy have proposed many different theories for how humans represent {``}concepts{''}. Multiple such theories are compatible with state-of-the-art NLP methods, and could in princi…

Philosophy