paper-with-me

홈 › Papers

QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning

2022-05-06 · Findings (NAACL) 2022 7 · Zechen Li, Anders Søgaard

Synthetic datasets have successfully been used to probe visual question-answering datasets for their reasoning abilities. CLEVR (johnson2017clevr), for example, tests a range of visual reasoning abilities. The questions in CLEVR focus on comparisons of shapes, colors, and sizes, numerical reasoning, and existence claims. This paper introduces a minimally biased, diagnostic visual question-answering dataset, QLEVR, that goes beyond existential and numerical quantification and focus on more complex quantifiers and their combinations, e.g., asking whether there are more than two red balls that are smaller than at least three blue balls in an image. We describe how the dataset was created and present a first evaluation of state-of-the-art visual question-answering models, showing that QLEVR presents a formidable challenge to our current models. Code and Dataset are available at https://github.com/zechenli03/QLEVR

📄 PDF Abstract BibTeX arXiv:2205.03075

Code (1)

zechenli03/qlevr 공식 구현

Tasks

DiagnosticQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Similar Papers 제목 키워드 기반

How Do People Quantify Naturally: Evidence from Mandarin Picture Description

2026-02-10 · Yayun Zhang, Guanyi Chen, Fahime Same, Saad Mahamood 외 arxiv

Quantification is a fundamental component of everyday language use, yet little is known about how speakers decide whether and how to quantify in naturalistic production. We investigate quantification in Mandarin Chinese …

Quantificational features in distributional word representations

2016-08-01 · SEMEVAL 2016 8 · Tal Linzen, Emmanuel Dupoux, Benjamin Spector

CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning

2016-12-20 · CVPR 2017 7 · Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei 외

When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover shortcomings. Existing benchmarks for visual question an…

DiagnosticQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)+1

Physics-informed Shadowgraph Network: An End-to-end Density Field Reconstruction Method

2024-10-26 · Xutun Wang, Yuchen Zhang, Zidong Li, Haocheng Wen 외

This study presents a novel approach for quantificationally reconstructing density fields from shadowgraph images using physics-informed neural networks

CMATH: Can Your Language Model Pass Chinese Elementary School Math Test?

2023-06-29 · Tianwen Wei, Jian Luan, Wei Liu, Shuang Dong 외

We present the Chinese Elementary School Math Word Problems (CMATH) dataset, comprising 1.7k elementary school-level math word problems with detailed annotations, source from actual Chinese workbooks and exams. This data…

Language ModelingLanguage ModellingMathMath Word Problem Solving