paper-with-me

Papers

Deep Bayesian Network for Visual Question Generation

2020-01-23 · Badri N. Patro, Vinod K. Kurmi, Sandeep Kumar, Vinay P. Namboodiri

Generating natural questions from an image is a semantic task that requires using vision and language modalities to learn multimodal representations. Images can have multiple visual and language cues such as places, captions, and tags. In this paper, we propose a principled deep Bayesian learning framework that combines these cues to produce natural questions. We observe that with the addition of more cues and by minimizing uncertainty in the among cues, the Bayesian network becomes more confident. We propose a Minimizing Uncertainty of Mixture of Cues (MUMC), that minimizes uncertainty present in a mixture of cues experts for generating probabilistic questions. This is a Bayesian framework and the results show a remarkable similarity to natural questions as validated by a human study. We observe that with the addition of more cues and by minimizing uncertainty among the cues, the Bayesian framework becomes more confident. Ablation studies of our model indicate that a subset of cues is inferior at this task and hence the principled fusion of cues is preferred. Further, we observe that the proposed approach substantially improves over state-of-the-art benchmarks on the quantitative metrics (BLEU-n, METEOR, ROUGE, and CIDEr). Here we provide project link for Deep Bayesian VQG \url{https://delta-lab-iitk.github.io/BVQG/}

📄 PDF Abstract BibTeX arXiv:2001.08779

Code (0)

등록된 구현이 없습니다.

Tasks

Natural QuestionsQuestion GenerationQuestion-Generation

Similar Papers 제목 키워드 기반

BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering

2026-04-24 · Jinghong Chen, Jingbiao Mei, Guangyu Yang, Bill Byrne arxiv

A common approach to question answering with retrieval-augmented generation (RAG) is to concatenate documents into a single context and pass it to a language model to generate an answer. While simple, this strategy can o…

Visual Question Answering

Deep Bayesian Active Learning for Multiple Correct Outputs

2019-12-02 · Khaled Jedoui, Ranjay Krishna, Michael Bernstein, Li Fei-Fei

Typical active learning strategies are designed for tasks, such as classification, with the assumption that the output space is mutually exclusive. The assumption that these tasks always have exactly one correct answer h…

Active LearningAnswer GenerationImage CaptioningQuestion-Answer-Generation+3

Answer-Type Prediction for Visual Question Answering

2016-06-01 · CVPR 2016 6 · Kushal Kafle, Christopher Kanan

Recently, algorithms for object recognition and related tasks have become sufficiently proficient that new vision tasks can now be pursued. In this paper, we build a system capable of answering open-ended text-based ques…

Object RecognitionPredictionQuestion AnsweringType prediction+3

What's to know? Uncertainty as a Guide to Asking Goal-oriented Questions

2018-12-16 · CVPR 2019 6 · Ehsan Abbasnejad, Qi Wu, Javen Shi, Anton Van Den Hengel

One of the core challenges in Visual Dialogue problems is asking the question that will provide the most useful information towards achieving the required objective. Encouraging an agent to ask the right questions is dif…

Visual Dialog

Comparing verbal, visual and combined explanations for Bayesian Network inferences

2025-11-21 · Erik P. Nyberg, Steven Mascaro, Ingrid Zukerman, Michael Wybrow 외 arxiv

Bayesian Networks (BNs) are an important tool for assisting probabilistic reasoning, but despite being considered transparent models, people have trouble understanding them. Further, current User Interfaces (UIs) still d…