paper-with-me

홈 › Papers

VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization

2023-11-01 · Suraj Jyothi Unni, Raha Moraffah, Huan Liu

Visual question answering (VQA) models are designed to demonstrate visual-textual reasoning capabilities. However, their real-world applicability is hindered by a lack of comprehensive benchmark datasets. Existing domain generalization datasets for VQA exhibit a unilateral focus on textual shifts while VQA being a multi-modal task contains shifts across both visual and textual domains. We propose VQA-GEN, the first ever multi-modal benchmark dataset for distribution shift generated through a shift induced pipeline. Experiments demonstrate VQA-GEN dataset exposes the vulnerability of existing methods to joint multi-modal distribution shifts. validating that comprehensive multi-modal shifts are critical for robust VQA generalization. Models trained on VQA-GEN exhibit improved cross-domain and in-domain performance, confirming the value of VQA-GEN. Further, we analyze the importance of each shift technique of our pipeline contributing to the generalization of the model.

📄 PDF Abstract BibTeX arXiv:2311.00807

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

An Empirical Study on the Language Modal in Visual Question Answering

2023-05-17 · Daowan Peng, Wei Wei, Xian-Ling Mao, Yuanyuan Fu 외

Generalization beyond in-domain experience to out-of-distribution data is of paramount significance in the AI domain. Of late, state-of-the-art Visual Question Answering (VQA) models have shown impressive performance on …

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation

2026-05-05 · Quanxing Xu, Ling Zhou, Xian Zhong, Xiaohua Huang 외 arxiv

With advances in multimodal research and deep learning, Multimodal Large Language Models (MLLMs) have emerged as a powerful paradigm for a wide range of multimodal tasks. As a core problem in vision-language research, Vi…

Visual Question Answering

Encoder Adaptation of Dense Passage Retrieval for Open-Domain Question Answering

2021-10-04 · Minghan Li, Jimmy Lin

One key feature of dense passage retrievers (DPR) is the use of separate question and passage encoder in a bi-encoder design. Previous work on generalization of DPR mainly focus on testing both encoders in tandem on out-…

Domain AdaptationOpen-Domain Question AnsweringPassage RetrievalQuestion Answering+1

NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding

2025-04-12 · Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh 외

Understanding and reasoning over academic handwritten notes remains a challenge in document AI, particularly for mathematical equations, diagrams, and scientific notations. Existing visual question answering (VQA) benchm…

BenchmarkingDocument AIdocument understandingMultimodal Reasoning+6

Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning

2022-12-01 · CVPR 2023 1 · Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski 외

Visual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple factors of variation are intertwined, ma…

Domain GeneralizationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)+1