Visual Reasoning with Natural Language
Natural language provides a widely accessible and expressive interface for robotic agents. To understand language in complex environments, agents must reason about the full range of language inputs and their correspondence to the world. Such reasoning over language and vision is an open problem that is receiving increasing attention. While existing data sets focus on visual diversity, they do not display the full range of natural language expressions, such as counting, set reasoning, and comparisons. We propose a simple task for natural language visual reasoning, where images are paired with descriptive statements. The task is to predict if a statement is true for the given scene. This abstract describes our existing synthetic images corpus and our current work on collecting real vision data.
Code (0)
등록된 구현이 없습니다.
Tasks
DescriptiveDiversityVisual ReasoningSimilar Papers 제목 키워드 기반
A Corpus for Reasoning About Natural Language Grounded in Photographs
We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sente…
DiversityVisual ReasoningDeepVIS: Bridging Natural Language and Data Visualization Through Step-wise Reasoning
Although data visualization is powerful for revealing patterns and communicating insights, creating effective visualizations requires familiarity with authoring tools and often disrupts the analysis flow. While large lan…
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs
Natural language rationales could provide intuitive, higher-level explanations that are easily understandable by humans, complementing the more broadly studied lower-level explanations based on gradients or attention wei…
Language ModelingLanguage ModellingNatural Language InferenceObject Recognition+5A Corpus of Natural Language for Visual Reasoning
We present a new visual reasoning language dataset, containing 92,244 pairs of examples of natural statements grounded in synthetic images with 3,962 unique sentences. We describe a method of crowdsourcing linguistically…
Question AnsweringVisual Question Answering (VQA)Visual ReasoningQ-Tacit: Image Quality Assessment via Latent Visual Reasoning
Vision-Language Model (VLM)-based image quality assessment (IQA) has been significantly advanced by incorporating Chain-of-Thought (CoT) reasoning. Recent work has refined image quality reasoning by applying reinforcemen…
Image Quality AssessmentReinforcement LearningVisual Reasoning