Multimodal Differential Network for Visual Question Generation
Generating natural questions from an image is a semantic task that requires using visual and language modality to learn multimodal representations. Images can have multiple visual and language contexts that are relevant for generating questions namely places, captions, and tags. In this paper, we propose the use of exemplars for obtaining the relevant context. We obtain this by using a Multimodal Differential Network to produce natural and engaging questions. The generated questions show a remarkable similarity to the natural questions as validated by a human study. Further, we observe that the proposed approach substantially improves over state-of-the-art benchmarks on the quantitative metrics (BLEU, METEOR, ROUGE, and CIDEr).
Code (1)
Tasks
Natural QuestionsQuestion GenerationQuestion-GenerationSimilar Papers 제목 키워드 기반
Multimodal Differential Network for Visual Question Generation
Generating natural questions from an image is a semantic task that requires using visual and language modality to learn multimodal representations. Images can have multiple visual and language contexts that are relevant …
Image CaptioningNatural QuestionsQuestion AnsweringQuestion Generation+2MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants
In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken in…
intent-classificationIntent ClassificationQuestion AnsweringQuestion Generation+3Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question Generation
Multiple-choice questions (MCQs) are important in enhancing concept learning and student engagement for educational purposes. Despite the multimodal nature of educational content, current methods focus mainly on text-bas…
Distractor GenerationMultiple-choiceQuestion GenerationQuestion-Generation+1DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introdu…
Spatial ReasoningVisual ReasoningVisual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
Retrieval-Augmented Generation (RAG) is a popular approach for enhancing Large Language Models (LLMs) by addressing their limitations in verifying facts and answering knowledge-intensive questions. As the research in LLM…
BenchmarkingImage RetrievalQuestion AnsweringRAG+2