Recent, rapid advancement in visual question answering architecture: a review
Understanding visual question answering is going to be crucial for numerous human activities. However, it presents major challenges at the heart of the artificial intelligence endeavor. This paper presents an update on the rapid advancements in visual question answering using images that have occurred in the last couple of years. Tremendous growth in research on improving visual question answering system architecture has been published recently, showing the importance of multimodal architectures. Several points on the benefits of visual question answering are mentioned in the review paper by Manmadhan et al. (2020), on which the present article builds, including subsequent updates in the field.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression
Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compa…
Question AnsweringThe Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering
Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about…
Common Sense ReasoningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations
We propose a novel framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. Traditional evaluation methods often rely on human judgment, which is co…
Chart Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)ORD: Object Relationship Discovery for Visual Dialogue Generation
With the rapid advancement of image captioning and visual question answering at single-round level, the question of how to generate multi-round dialogue about visual content has not yet been well explored.Existing visual…
Dialogue GenerationGraph AttentionImage CaptioningObject+4Targeted Visual Prompting for Medical Visual Question Answering
With growing interest in recent years, medical visual question answering (Med-VQA) has rapidly evolved, with multimodal large language models (MLLMs) emerging as an alternative to classical model architectures. Specifica…
Medical Visual Question AnsweringQuestion AnsweringVisual PromptingVisual Question Answering+1