paper-with-me

Papers

Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering

2023-06-08 · Param Ahir, Dr. Hiteishi Diwanji

Visual question answering (VQA) is a Multidisciplinary research problem that pursued through practices of natural language processing and computer vision. Visual question answering automatically answers natural language questions according to the content of an image. Some testing questions require external knowledge to derive a solution. Such knowledge-based VQA uses various methods to retrieve features of image and text, and combine them to generate the answer. To generate knowledgebased answers either question dependent or image dependent knowledge retrieval methods are used. If knowledge about all the objects in the image is derived, then not all knowledge is relevant to the question. On other side only question related knowledge may lead to incorrect answers and over trained model that answers question that is irrelevant to image. Our proposed method takes image attributes and question features as input for knowledge derivation module and retrieves only question relevant knowledge about image objects which can provide accurate answers.

📄 PDF Abstract BibTeX arXiv:2306.04938

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Personalized Questions, Answers and Grammars: Aiding the Search for Relevant Web Information

2017-09-01 · WS 2017 9 · Marta Gatius

This work proposes an organization of knowledge to facilitate the generation of personalized questions, answers and grammars from web documents. To reduce the human effort needed in the generation of the linguistic resou…

Text Generation

OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

2019-05-31 · CVPR 2019 6 · Kenneth Marino, Mohammad Rastegari, Ali Farhadi, Roozbeh Mottaghi

Visual Question Answering (VQA) in its ideal form lets us study reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most VQA benchmarks to date are f…

object-detectionObject DetectionQuestion AnsweringScene Understanding+2

GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering

2024-02-04 · Ziyu Ma, Shutao Li, Bin Sun, Jianfei Cai 외

Knowledge-based visual question answering (VQA) requires world knowledge beyond the image for accurate answer. Recently, instead of extra knowledge bases, a large language model (LLM) like GPT-3 is activated as an implic…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

Complex Event Detection via Multi-source Video Attributes

2013-06-01 · CVPR 2013 6 · Zhigang Ma, Yi Yang, Zhongwen Xu, Shuicheng Yan 외

Complex events essentially include human, scenes, objects and actions that can be summarized by visual attributes, so leveraging relevant attributes properly could be helpful for event detection. Many works have exploite…

Event Detection

GRADE: Quantifying Sample Diversity in Text-to-Image Models

2024-10-29 · Royi Rassin, Aviv Slobodkin, Shauli Ravfogel, Yanai Elazar 외

Text-to-image (T2I) models are remarkable at generating realistic images based on textual descriptions. However, textual prompts are inherently underspecified: they do not specify all possible attributes of the required …

AttributeDiversityQuestion AnsweringVisual Question Answering+1