paper-with-me

홈 › Papers

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering

2025-01-13 · Anupam Pandey, Deepjyoti Bodo, Arpan Phukan, Asif Ekbal

Visual Question Answering (VQA) is an interdisciplinary field that bridges the gap between computer vision (CV) and natural language processing(NLP), enabling Artificial Intelligence(AI) systems to answer questions about images. Since its inception in 2015, VQA has rapidly evolved, driven by advances in deep learning, attention mechanisms, and transformer-based models. This survey traces the journey of VQA from its early days, through major breakthroughs, such as attention mechanisms, compositional reasoning, and the rise of vision-language pre-training methods. We highlight key models, datasets, and techniques that shaped the development of VQA systems, emphasizing the pivotal role of transformer architectures and multimodal pre-training in driving recent progress. Additionally, we explore specialized applications of VQA in domains like healthcare and discuss ongoing challenges, such as dataset bias, model interpretability, and the need for common-sense reasoning. Lastly, we discuss the emerging trends in large multimodal language models and the integration of external knowledge, offering insights into the future directions of VQA. This paper aims to provide a comprehensive overview of the evolution of VQA, highlighting both its current state and potential advancements.

📄 PDF Abstract BibTeX arXiv:2501.07109

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Sketch & Paint: Stroke-by-Stroke Evolution of Visual Artworks

2025-02-27 · Jeripothula Prudviraj, Vikram Jamwal

Understanding the stroke-based evolution of visual artworks is useful for advancing artwork learning, appreciation, and interactive display. While the stroke sequence of renowned artworks remains largely unknown, formula…

Clustering

JourneyDB: A Benchmark for Generative Image Understanding

2023-07-03 · NeurIPS 2023 11

While recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehend generated images remains uncertain. Sy…

Image CaptioningImage ComprehensionQuestion AnsweringVisual Question Answering

JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images

2024-09-19 · Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani AlOmari 외

Existing vision-language understanding benchmarks largely consist of images of objects in their usual contexts. As a consequence, recent multimodal large language models can perform well with only a shallow visual unders…

HallucinationImage CaptioningMultimodal ReasoningVisual Question Answering (VQA)+1

Data Quality Awareness: A Journey from Traditional Data Management to Data Science Systems

2024-11-05 · Sijie Dong, Soror Sahri, Themis Palpanas

Artificial intelligence (AI) has transformed various fields, significantly impacting our daily lives. A major factor in AI success is high-quality data. In this paper, we present a comprehensive review of the evolution o…

Management

Large Language Models for User Interest Journeys

2023-05-24 · Konstantina Christakopoulou, Alberto Lalama, Cj Adams, Iris Qu 외

Large language models (LLMs) have shown impressive capabilities in natural language understanding and generation. Their potential for deeper user understanding and improved personalized user experience on recommendation …

Natural Language UnderstandingRecommendation Systems