paper-with-me

Papers

Towards Explainable 3D Grounded Visual Question Answering: A New Benchmark and Strong Baseline

2022-09-24 · Lichen Zhao, Daigang Cai, Jing Zhang, Lu Sheng, Dong Xu, Rui Zheng, Yinjie Zhao, Lipeng Wang, Xibo Fan

Recently, 3D vision-and-language tasks have attracted increasing research interest. Compared to other vision-and-language tasks, the 3D visual question answering (VQA) task is less exploited and is more susceptible to language priors and co-reference ambiguity. Meanwhile, a couple of recently proposed 3D VQA datasets do not well support 3D VQA task due to their limited scale and annotation methods. In this work, we formally define and address a 3D grounded VQA task by collecting a new 3D VQA dataset, referred to as FE-3DGQA, with diverse and relatively free-form question-answer pairs, as well as dense and completely grounded bounding box annotations. To achieve more explainable answers, we labelled the objects appeared in the complex QA pairs with different semantic types, including answer-grounded objects (both appeared and not appeared in the questions), and contextual objects for answer-grounded objects. We also propose a new 3D VQA framework to effectively predict the completely visually grounded and explainable answer. Extensive experiments verify that our newly collected benchmark datasets can be effectively used to evaluate various 3D VQA methods from different aspects and our newly proposed framework also achieves state-of-the-art performance on the new benchmark dataset. Both the newly collected dataset and our codes will be publicly available at http://github.com/zlccccc/3DGQA.

📄 PDF Abstract BibTeX arXiv:2209.12028

Code (1)

zlccccc/3dgqa 공식 구현 pytorch

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions

2026-05-07 · Kjetil Indrehus, Adrian Duric, Changkyu Choi, Ali Ramezani-Kebrya arxiv

Document Visual Question Answering (DocVQA) requires vision-language models to reason not only about what information in a document is relevant to a question, but also where the answer is grounded on the page. Existing D…

Visual Question Answering

VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning

2025-11-01 · Dang H. Nguyen, Hieu H. Pham, Hao T. Nguyen, Hieu H. Pham arxiv

We present VinDr-CXR-VQA, a large-scale chest X-ray dataset for explainable Medical Visual Question Answering (Med-VQA) with spatial grounding. The dataset contains 17,597 question-answer pairs across 4,394 images, each …

Visual Question AnsweringMulti-Task Learning

ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports

2025-11-20 · Sherine George, Nithish Saji arxiv

We present ESGBench, a benchmark dataset and evaluation framework designed to assess explainable ESG question answering systems using corporate sustainability reports. The benchmark consists of domain-grounded questions …

Question Answering

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

2025-04-03 · Hao Li, Hao Fei, Zechao Hu, Zhengwei Yang 외

Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model's social intelligence level. While impressive multiple-choice question(MCQ) accuracy is achieved by current solutio…

Multiple-choice

Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

2025-07-10 · Haochen Wang, Xiangtai Li, Zilong Huang, Anran Wang 외 arxiv

Models like OpenAI-o3 pioneer visual grounded reasoning by dynamically referencing visual regions, just like human "thinking with images". However, no benchmark exists to evaluate these capabilities holistically. To brid…

Reinforcement LearningObject Localization