Diversity and Consistency: Exploring Visual Question-Answer Pair Generation
Although showing promising values to downstream applications, generating question and answer together is under-explored. In this paper, we introduce a novel task that targets question-answer pair generation from visual images. It requires not only generating diverse question-answer pairs but also keeping the consistency of them. We study different generation paradigms for this task and propose three models: the pipeline model, the joint model, and the sequential model. We integrate variational inference into these models to achieve diversity and consistency. We also propose region representation scaling and attention alignment to improve the consistency further. We finally devise an evaluator as a quantitative metric for consistency. We validate our approach on two benchmarks, VQA2.0 and Visual-7w, by automatically and manually evaluating diversity and consistency. Experimental results show the effectiveness of our models: they can generate diverse or consistent pairs. Moreover, this task can be used to improve visual question generation and visual question answering.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityQuestion AnsweringQuestion GenerationQuestion-GenerationVariational InferenceVisual Question AnsweringVisual Question Answering (VQA)Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison
Visual Question Answering (VQA) has emerged as a pivotal task in the intersection of computer vision and natural language processing, requiring models to understand and reason about visual content in response to natural …
DiversityLanguage ModelingLanguage ModellingMultimodal Reasoning+3SimVQA: Exploring Simulated Environments for Visual Question Answering
Existing work on VQA explores data augmentation to achieve better generalization by perturbing the images in the dataset or modifying the existing questions and answers. While these methods exhibit good performance, the …
Data AugmentationDiversityQuestion AnsweringVisual Question Answering+1Diversify Question Generation with Retrieval-Augmented Style Transfer
Given a textual passage and an answer, humans are able to ask questions with various expressions, but this ability is still challenging for most question generation (QG) systems. Existing solutions mainly focus on the in…
DiversityQuestion AnsweringQuestion GenerationQuestion-Generation+3VQA Therapy: Exploring Answer Differences by Visually Grounding Answers
Visual question answering is a task of predicting the answer to a question about an image. Given that different people can provide different answers to a visual question, we aim to better understand why with answer groun…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets
Medical Visual Question Answering (MVQA) systems can interpret medical images in response to natural language queries. However, linguistic variability in question phrasing often undermines the consistency of these system…
DiversityMedical Visual Question AnsweringNatural Language QueriesQuestion Answering+3