Survey of Visual Question Answering: Datasets and Techniques
Visual question answering (or VQA) is a new and exciting problem that combines natural language processing and computer vision techniques. We present a survey of the various datasets and models that have been used to tackle this task. The first part of the survey details the various datasets for VQA and compares them along some common factors. The second part of this survey details the different approaches for VQA, classified into four types: non-deep learning models, deep learning models without attention, deep learning models with attention, and other models which do not fit into the first three. Finally, we compare the performances of these approaches and provide some directions for future work.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningQuestion AnsweringSurveyVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Recent Advances in Video Question Answering: A Review of Datasets and Methods
Video Question Answering (VQA) is a recent emerging challenging task in the field of Computer Vision. Several visual information retrieval techniques like Video Captioning/Description and Video-guided Machine Translation…
Information RetrievalMachine TranslationQuestion AnsweringRetrieval+6Visual Question Answering: A Survey on Techniques and Common Trends in Recent Literature
Visual Question Answering (VQA) is an emerging area of interest for researches, being a recent problem in natural language processing and image prediction. In this area, an algorithm needs to answer questions about certa…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Medical Visual Question Answering: A Survey
Medical Visual Question Answering~(VQA) is a combination of medical artificial intelligence and popular VQA challenges. Given a medical image and a clinically relevant question in natural language, the medical VQA system…
Medical Visual Question AnsweringQuestion AnsweringSurveyVisual Question Answering+1Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey
Visual Question Answering (VQA) is a challenge task that combines natural language processing and computer vision techniques and gradually becomes a benchmark test task in multimodal large language models (MLLMs). The go…
Natural Language UnderstandingQuestion AnsweringSurveyVisual Question Answering+1Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
Creating engaging narratives from visual data is crucial for automated digital media consumption, assistive technologies, and interactive entertainment. This survey covers methodologies used in the generation of these na…
Question AnsweringStory GenerationSurveyVideo Captioning+1