paper-with-me

홈 › Papers

Recent, rapid advancement in visual question answering architecture: a review

2022-03-02 · Venkat Kodali, Daniel Berleant

Understanding visual question answering is going to be crucial for numerous human activities. However, it presents major challenges at the heart of the artificial intelligence endeavor. This paper presents an update on the rapid advancements in visual question answering using images that have occurred in the last couple of years. Tremendous growth in research on improving visual question answering system architecture has been published recently, showing the importance of multimodal architectures. Several points on the benefits of visual question answering are mentioned in the review paper by Manmadhan et al. (2020), on which the present article builds, including subsequent updates in the field.

📄 PDF Abstract BibTeX arXiv:2203.01322

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Zero-shot 3D Question Answering via Voxel-based Dynamic Token Compression

2025-01-01 · CVPR 2025 1 · Hsiang-Wei Huang, Fu-Chen Chen, Wenhao Chai, Che-Chun Su 외

Recent advancements in 3D Large Multi-modal Models (3D-LMMs) have driven significant progress in 3D question answering. However, recent multi-frame Vision-Language Models (VLMs) demonstrate superior performance compa…

Question Answering

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

2025-01-13 · Anupam Pandey, Deepjyoti Bodo, Arpan Phukan, Asif Ekbal

Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about…

Common Sense ReasoningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations

2024-09-27 · James Ford, Xingmeng Zhao, Dan Schumacher, Anthony Rios

We propose a novel framework that leverages Visual Question Answering (VQA) models to automate the evaluation of LLM-generated data visualizations. Traditional evaluation methods often rely on human judgment, which is co…

Chart Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

ORD: Object Relationship Discovery for Visual Dialogue Generation

2020-06-15 · Ziwei Wang, Zi Huang, Yadan Luo, Huimin Lu

With the rapid advancement of image captioning and visual question answering at single-round level, the question of how to generate multi-round dialogue about visual content has not yet been well explored.Existing visual…

Dialogue GenerationGraph AttentionImage CaptioningObject+4

Targeted Visual Prompting for Medical Visual Question Answering

2024-08-06 · Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman

With growing interest in recent years, medical visual question answering (Med-VQA) has rapidly evolved, with multimodal large language models (MLLMs) emerging as an alternative to classical model architectures. Specifica…

Medical Visual Question AnsweringQuestion AnsweringVisual PromptingVisual Question Answering+1