paper-with-me

Papers

Visual Question Answering Dataset for Bilingual Image Understanding: A Study of Cross-Lingual Transfer Using Attention Maps

2018-08-01 · COLING 2018 8 · Nobuyuki Shimizu, Na Rong, Takashi Miyazaki

Visual question answering (VQA) is a challenging task that requires a computer system to understand both a question and an image. While there is much research on VQA in English, there is a lack of datasets for other languages, and English annotation is not directly applicable in those languages. To deal with this, we have created a Japanese VQA dataset by using crowdsourced annotation with images from the Visual Genome dataset. This is the first such dataset in Japanese. As another contribution, we propose a cross-lingual method for making use of English annotation to improve a Japanese VQA system. The proposed method is based on a popular VQA method that uses an attention mechanism. We use attention maps generated from English questions to help improve the Japanese VQA task. The proposed method experimentally performed better than simply using a monolingual corpus, which demonstrates the effectiveness of using attention maps to transfer cross-lingual information.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferImage CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining

2024-01-12 · Minjun Kim, Seungwoo Song, Youhan Lee, Haneol Jang 외

The current research direction in generative models, such as the recently developed GPT4, aims to find relevant knowledge information for multimodal and multilingual inputs to provide answers. Under these research circum…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

On the General Value of Evidence, and Bilingual Scene-Text Visual Question Answering

2020-02-24 · CVPR 2020 6 · Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng 외

Visual Question Answering (VQA) methods have made incredible progress, but suffer from a failure to generalize. This is visible in the fact that they are vulnerable to learning coincidental correlations in the data rathe…

Question AnsweringReferring ExpressionVisual Question AnsweringVisual Question Answering (VQA)

Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat

2025-05-26 · Pusheng Xu, Xia Gong, Xiaolan Chen, Weiyi Zhang 외

Purpose: To develop a bilingual multimodal visual question answering (VQA) benchmark for evaluating VLMs in ophthalmology. Methods: Ophthalmic image posts and associated captions published between January 1, 2016, and De…

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

TextlessRAG: End-to-End Visual Document RAG by Speech Without Text

2025-09-09 · Peijin Xie, Shun Qian, Bingquan Liu, Dexin Wang 외 arxiv

Document images encapsulate a wealth of knowledge, while the portability of spoken queries enables broader and flexible application scenarios. Yet, no prior work has explored knowledge base question answering over visual…

Knowledge Base Question Answering

Multilingual Hematology Visual Question Answering Dataset

2026-06-24 · Hajra Malik, Hafiza Tooba Aftab, Abdul Rehman, Mohsen Ali 외 arxiv

Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology …

Visual Question Answering