paper-with-me

홈 › Papers

ViVQA: Vietnamese Visual Question Answering

2021-11-01 · PACLIC 2021 11 · Khanh Quoc Tran, An Trong Nguyen, An Tran-Hoai Le, Kiet Van Nguyen
📄 PDF Abstract BibTeX

Code (2)

khanhtran0412/vivqa 공식 구현
hieunghia-pat/UIT-MCAN pytorch

Tasks

Question AnsweringVietnamese Visual Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

2026-03-10 · Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh 외 arxiv

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work …

Visual Question AnsweringRepresentation LearningMachine TranslationImage Captioning

OpenViVQA: Task, Dataset, and Multimodal Fusion Models for Visual Question Answering in Vietnamese

2023-05-07 · Nghia Hieu Nguyen, Duong T. D. Vo, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

In recent years, visual question answering (VQA) has attracted attention from the research community because of its highly potential applications (such as virtual assistance on intelligent cars, assistant devices for bli…

Information RetrievalQuestion AnsweringRetrievalVietnamese Multimodal Learning+4

Advancing Vietnamese Visual Question Answering with Transformer and Convolutional Integration

2024-07-30

Visual Question Answering (VQA) has recently emerged as a potential research domain, captivating the interest of many in the field of artificial intelligence and computer vision. Despite the prevalence of approaches in E…

Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese

2024-08-22 · Khang T. Doan, Bao G. Huynh, Dung T. Hoang, Thuc D. Pham 외

In this report, we introduce Vintern-1B, a reliable 1-billion-parameters multimodal large language model (MLLM) for Vietnamese language tasks. By integrating the Qwen2-0.5B-Instruct language model with the InternViT-300M…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

PAT: Parallel Attention Transformer for Visual Question Answering in Vietnamese

2023-07-17 · Nghia Hieu Nguyen, Kiet Van Nguyen

We present in this paper a novel scheme for multimodal learning named the Parallel Attention mechanism. In addition, to take into account the advantages of grammar and context in Vietnamese, we propose the Hierarchical L…

Question AnsweringVietnamese Visual Question AnsweringVisual Question Answering