paper-with-me

Papers

Multimodal Differential Network for Visual Question Generation

2018-08-12 · EMNLP 2018 10 · Badri N. Patro, Sandeep Kumar, Vinod K. Kurmi, Vinay P. Namboodiri

Generating natural questions from an image is a semantic task that requires using visual and language modality to learn multimodal representations. Images can have multiple visual and language contexts that are relevant for generating questions namely places, captions, and tags. In this paper, we propose the use of exemplars for obtaining the relevant context. We obtain this by using a Multimodal Differential Network to produce natural and engaging questions. The generated questions show a remarkable similarity to the natural questions as validated by a human study. Further, we observe that the proposed approach substantially improves over state-of-the-art benchmarks on the quantitative metrics (BLEU, METEOR, ROUGE, and CIDEr).

📄 PDF Abstract BibTeX arXiv:1808.03986

Code (1)

badripatro/MDN-VQG pytorch

Tasks

Natural QuestionsQuestion GenerationQuestion-Generation

Similar Papers 제목 키워드 기반

Multimodal Differential Network for Visual Question Generation

2018-10-01 · EMNLP 2018 10 · Badri Narayana Patro, S. Kumar, eep, Vinod Kumar Kurmi 외

Generating natural questions from an image is a semantic task that requires using visual and language modality to learn multimodal representations. Images can have multiple visual and language contexts that are relevant …

Image CaptioningNatural QuestionsQuestion AnsweringQuestion Generation+2

MMIU: Dataset for Visual Intent Understanding in Multimodal Assistants

2021-10-13 · Alkesh Patel, Joel Ruben Antony Moniz, Roman Nguyen, Nick Tzou 외

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken in…

intent-classificationIntent ClassificationQuestion AnsweringQuestion Generation+3

Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question Generation

2024-08-16 · Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 2024 8 · Haohao Luo, Yang Deng, Ying Shen, See-Kiong Ng 외

Multiple-choice questions (MCQs) are important in enhancing concept learning and student engagement for educational purposes. Despite the multimodal nature of educational content, current methods focus mainly on text-bas…

Distractor GenerationMultiple-choiceQuestion GenerationQuestion-Generation+1

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

2025-12-14 · Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao 외 arxiv

Multimodal Large Language Models have achieved impressive performance on a variety of vision-language tasks, yet their fine-grained visual perception and precise spatial reasoning remain limited. In this work, we introdu…

Spatial ReasoningVisual Reasoning

Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

2025-02-23 · Yin Wu, Quanyu Long, Jing Li, Jianfei Yu 외

Retrieval-Augmented Generation (RAG) is a popular approach for enhancing Large Language Models (LLMs) by addressing their limitations in verifying facts and answering knowledge-intensive questions. As the research in LLM…

BenchmarkingImage RetrievalQuestion AnsweringRAG+2