paper-with-me

홈 › Papers

A Corpus of Natural Language for Visual Reasoning

2017-07-01 · ACL 2017 7 · Alane Suhr, Mike Lewis, James Yeh, Yoav Artzi

We present a new visual reasoning language dataset, containing 92,244 pairs of examples of natural statements grounded in synthetic images with 3,962 unique sentences. We describe a method of crowdsourcing linguistically-diverse data, and present an analysis of our data. The data demonstrates a broad set of linguistic phenomena, requiring visual and set-theoretic reasoning. We experiment with various models, and show the data presents a strong challenge for future research.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Similar Papers 제목 키워드 기반

Visual Reasoning with Natural Language

2017-10-02 · Stephanie Zhou, Alane Suhr, Yoav Artzi

Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their corresponden…

DescriptiveDiversityVisual Reasoning

A Corpus for Reasoning About Natural Language Grounded in Photographs

2018-11-01 · ACL 2019 7 · Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang 외

We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sente…

DiversityVisual Reasoning

UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning

2023-05-24 · Ahmed Masry, Parsa Kavehzadeh, Xuan Long Do, Enamul Hoque 외

Charts are very popular for analyzing data, visualizing key insights and answering complex reasoning questions about data. To facilitate chart-based data analysis using natural language, several downstream tasks have bee…

Chart Question AnsweringChart UnderstandingDecoderQuestion Answering

e-SNLI-VE: Corrected Visual-Textual Entailment with Natural Language Explanations

2020-04-07 · Virginie Do, Oana-Maria Camburu, Zeynep Akata, Thomas Lukasiewicz

The recently proposed SNLI-VE corpus for recognising visual-textual entailment is a large, real-world dataset for fine-grained multimodal reasoning. However, the automatic way in which SNLI-VE has been assembled (via com…

Multimodal ReasoningNatural Language Inference

Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning

2025-07-22 · Ang Li, Charles Wang, Deqing Fu, Kaiyu Yue 외 arxiv

Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual CoT), is challenging due to: (1) poor off…

Reinforcement LearningMultimodal ReasoningVisual Reasoning