The Meaning of ``Most'' for Visual Question Answering Models
The correct interpretation of quantifier statements in the context of a visual scene requires non-trivial inference mechanisms. For the example of {``}most{''}, we discuss two strategies which rely on fundamentally different cognitive concepts. Our aim is to identify what strategy deep learning models for visual question answering learn when trained on such questions. To this end, we carefully design data to replicate experiments from psycholinguistics where the same question was investigated for humans. Focusing on the FiLM visual question answering model, our experiments indicate that a form of approximate number system emerges whose performance declines with more difficult scenes as predicted by Weber{'}s law. Moreover, we identify confounding factors, like spatial arrangement of the scene, which impede the effectiveness of this system.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
The meaning of "most" for visual question answering models
The correct interpretation of quantifier statements in the context of a visual scene requires non-trivial inference mechanisms. For the example of "most", we discuss two strategies which rely on fundamentally different c…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)'Just because you are right, doesn't mean I am wrong': Overcoming a Bottleneck in the Development and Evaluation of Open-Ended Visual Question Answering (VQA) Tasks
GQA~\citep{hudson2019gqa} is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the gro…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningMultiple Meta-model Quantifying for Medical Visual Question Answering
Transfer learning is an important step to extract meaningful features and overcome the data limitation in the medical Visual Question Answering (VQA) task. However, most of the existing medical VQA methods rely on extern…
Medical Visual Question AnsweringMeta-LearningQuestion AnsweringTransfer Learning+2`Just because you are right, doesn't mean I am wrong': Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks
GQA (CITATION) is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models on the GQA dataset do not match the ground-truth a…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningMultilingual Augmentation for Robust Visual Question Answering in Remote Sensing Images
Aiming at answering questions based on the content of remotely sensed images, visual question answering for remote sensing data (RSVQA) has attracted much attention nowadays. However, previous works in RSVQA have focused…
Contrastive LearningQuestion AnsweringVisual Question Answering