paper-with-me

홈 › Papers

Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark

2025-12-22 · Hieu Minh Nguyen, Tam Le-Thanh Dang, Kiet Van Nguyen arxiv

Understanding signboard text in natural scenes is essential for real-world applications of Visual Question Answering (VQA), yet remains underexplored, particularly in low-resource languages. We introduce ViSignVQA, the first large-scale Vietnamese dataset designed for signboard-oriented VQA, which comprises 10,762 images and 25,573 question-answer pairs. The dataset captures the diverse linguistic, cultural, and visual characteristics of Vietnamese signboards, including bilingual text, informal phrasing, and visual elements such as color and layout. To benchmark this task, we adapted state-of-the-art VQA models (e.g., BLIP-2, LaTr, PreSTU, and SaL) by integrating a Vietnamese OCR model (SwinTextSpotter) and a Vietnamese pretrained language model (ViT5). The experimental results highlight the significant role of the OCR-enhanced context, with F1-score improvements of up to 209% when the OCR text is appended to questions. Additionally, we propose a multi-agent VQA framework combining perception and reasoning agents with GPT-4, achieving 75.98% accuracy via majority voting. Our study presents the first large-scale multimodal dataset for Vietnamese signboard understanding. This underscores the importance of domain-specific resources in enhancing text-based VQA for low-resource languages. ViSignVQA serves as a benchmark capturing real-world scene text characteristics and supporting the development and evaluation of OCR-integrated VQA models in Vietnamese.

📄 PDF Abstract BibTeX arXiv:2512.22218

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Making the V in Text-VQA Matter

2023-08-01 · Shamanthak Hegde, Soumya Jahagirdar, Shankar Gangisetty

Text-based VQA aims at answering questions by reading the text present in the images. It requires a large amount of scene-text relationship understanding compared to the VQA task. Recent studies have shown that the quest…

Optical Character Recognition (OCR)TextVQAVisual Question Answering (VQA)

VizWiz Grand Challenge: Answering Visual Questions from Blind People

2018-02-22 · CVPR 2018 6 · Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo 외

The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA d…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Goal-Oriented Semantic Communication for Wireless Visual Question Answering

2024-11-03 · Sige Liu, Nan Li, Yansha Deng, Tony Q. S. Quek

The rapid progress of artificial intelligence (AI) and computer vision (CV) has facilitated the development of computation-intensive applications like Visual Question Answering (VQA), which integrates visual perception a…

Edge-computingQuestion AnsweringSemantic CommunicationVisual Question Answering+1

REXUP: I REason, I EXtract, I UPdate with Structured Compositional Reasoning for Visual Question Answering

2020-07-27 · Siwen Luo, Soyeon Caren Han, Kaiyuan Sun, Josiah Poon

Visual question answering (VQA) is a challenging multi-modal task that requires not only the semantic understanding of both images and questions, but also the sound perception of a step-by-step reasoning process that wou…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Automatic Signboard Detection and Localization in Densely Populated Developing Cities

2020-03-04 · Md. Sadrul Islam Toaha, Sakib Bin Asad, Chowdhury Rafeed Rahman, S. M. Shahriar Haque 외

Most city establishments of developing cities are digitally unlabeled because of the lack of automatic annotation systems. Hence location and trajectory services such as Google Maps, Uber etc remain underutilized in such…

Information RetrievalNovel Object Detectionobject-detectionObject Detection+1