paper-with-me

Papers

Diversity and Consistency: Exploring Visual Question-Answer Pair Generation

2021-11-01 · Findings (EMNLP) 2021 11 · Sen yang, Qingyu Zhou, Dawei Feng, Yang Liu, Chao Li, Yunbo Cao, Dongsheng Li

Although showing promising values to downstream applications, generating question and answer together is under-explored. In this paper, we introduce a novel task that targets question-answer pair generation from visual images. It requires not only generating diverse question-answer pairs but also keeping the consistency of them. We study different generation paradigms for this task and propose three models: the pipeline model, the joint model, and the sequential model. We integrate variational inference into these models to achieve diversity and consistency. We also propose region representation scaling and attention alignment to improve the consistency further. We finally devise an evaluator as a quantitative metric for consistency. We validate our approach on two benchmarks, VQA2.0 and Visual-7w, by automatically and manually evaluating diversity and consistency. Experimental results show the effectiveness of our models: they can generate diverse or consistent pairs. Moreover, this task can be used to improve visual question generation and visual question answering.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityQuestion AnsweringQuestion GenerationQuestion-GenerationVariational InferenceVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Variational Inference 설명 없음

Similar Papers 제목 키워드 기반

Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison

2025-02-20 · Aiswarya Baby, Tintu Thankom Koshy

Visual Question Answering (VQA) has emerged as a pivotal task in the intersection of computer vision and natural language processing, requiring models to understand and reason about visual content in response to natural …

DiversityLanguage ModelingLanguage ModellingMultimodal Reasoning+3

SimVQA: Exploring Simulated Environments for Visual Question Answering

2022-03-31 · CVPR 2022 1 · Paola Cascante-Bonilla, Hui Wu, Letao Wang, Rogerio Feris 외

Existing work on VQA explores data augmentation to achieve better generalization by perturbing the images in the dataset or modifying the existing questions and answers. While these methods exhibit good performance, the …

Data AugmentationDiversityQuestion AnsweringVisual Question Answering+1

Diversify Question Generation with Retrieval-Augmented Style Transfer

2023-10-23 · Qi Gou, Zehua Xia, Bowen Yu, Haiyang Yu 외

Given a textual passage and an answer, humans are able to ask questions with various expressions, but this ability is still challenging for most question generation (QG) systems. Existing solutions mainly focus on the in…

DiversityQuestion AnsweringQuestion GenerationQuestion-Generation+3

VQA Therapy: Exploring Answer Differences by Visually Grounding Answers

2023-08-21 · ICCV 2023 1 · Chongyan Chen, Samreen Anjum, Danna Gurari

Visual question answering is a task of predicting the answer to a question about an image. Given that different people can provide different answers to a visual question, we aim to better understand why with answer groun…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets

2025-04-16 · Yongpei Ma, Pengyu Wang, Adam Dunn, Usman Naseem 외

Medical Visual Question Answering (MVQA) systems can interpret medical images in response to natural language queries. However, linguistic variability in question phrasing often undermines the consistency of these system…

DiversityMedical Visual Question AnsweringNatural Language QueriesQuestion Answering+3