paper-with-me

홈 › Papers

The Meaning of ``Most'' for Visual Question Answering Models

2019-08-01 · WS 2019 8 · Alex Kuhnle, er, Ann Copestake

The correct interpretation of quantifier statements in the context of a visual scene requires non-trivial inference mechanisms. For the example of {``}most{''}, we discuss two strategies which rely on fundamentally different cognitive concepts. Our aim is to identify what strategy deep learning models for visual question answering learn when trained on such questions. To this end, we carefully design data to replicate experiments from psycholinguistics where the same question was investigated for humans. Focusing on the FiLM visual question answering model, our experiments indicate that a form of approximate number system emerges whose performance declines with more difficult scenes as predicted by Weber{'}s law. Moreover, we identify confounding factors, like spatial arrangement of the scene, which impede the effectiveness of this system.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

The meaning of "most" for visual question answering models

2018-12-31 · Alexander Kuhnle, Ann Copestake

The correct interpretation of quantifier statements in the context of a visual scene requires non-trivial inference mechanisms. For the example of "most", we discuss two strategies which rely on fundamentally different c…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

'Just because you are right, doesn't mean I am wrong': Overcoming a Bottleneck in the Development and Evaluation of Open-Ended Visual Question Answering (VQA) Tasks

2021-03-28 · Man Luo, Shailaja Keyur Sampat, Riley Tallman, Yankai Zeng 외

GQA~\citep{hudson2019gqa} is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the gro…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Multiple Meta-model Quantifying for Medical Visual Question Answering

2021-05-19 · Tuong Do, Binh X. Nguyen, Erman Tjiputra, Minh Tran 외

Transfer learning is an important step to extract meaningful features and overcome the data limitation in the medical Visual Question Answering (VQA) task. However, most of the existing medical VQA methods rely on extern…

Medical Visual Question AnsweringMeta-LearningQuestion AnsweringTransfer Learning+2

`Just because you are right, doesn't mean I am wrong': Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks

2021-04-01 · EACL 2021 2 · Man Luo, Shailaja Keyur Sampat, Riley Tallman, Yankai Zeng 외

GQA (CITATION) is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the ground-truth a…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Multilingual Augmentation for Robust Visual Question Answering in Remote Sensing Images

2023-04-07 · Zhenghang Yuan, Lichao Mou, Xiao Xiang Zhu

Aiming at answering questions based on the content of remotely sensed images, visual question answering for remote sensing data (RSVQA) has attracted much attention nowadays. However, previous works in RSVQA have focused…

Contrastive LearningQuestion AnsweringVisual Question Answering