paper-with-me

홈 › Papers

VIMQA: A Vietnamese Dataset for Advanced Reasoning and Explainable Multi-hop Question Answering

2022-06-01 · LREC 2022 6 · Khang Le, Hien Nguyen, Tung Le Thanh, Minh Nguyen

Vietnamese is the native language of over 98 million people in the world. However, existing Vietnamese Question Answering (QA) datasets do not explore the model’s ability to perform advanced reasoning and provide evidence to explain the answer. We introduce VIMQA, a new Vietnamese dataset with over 10,000 Wikipedia-based multi-hop question-answer pairs. The dataset is human-generated and has four main features: (1) The questions require advanced reasoning over multiple paragraphs. (2) Sentence-level supporting facts are provided, enabling the QA model to reason and explain the answer. (3) The dataset offers various types of reasoning to test the model’s ability to reason and extract relevant proof. (4) The dataset is in Vietnamese, a low-resource language. We also conduct experiments on our dataset using state-of-the-art Multilingual single-hop and multi-hop QA methods. The results suggest that our dataset is challenging for existing methods, and there is room for improvement in Vietnamese QA systems. In addition, we propose a general process for data creation and publish a framework for creating multilingual multi-hop QA datasets. The dataset and framework are publicly available to encourage further research in Vietnamese QA systems.

📄 PDF Abstract BibTeX

Code (1)

vimqa/vimqa 공식 구현

Tasks

Multi-hop Question AnsweringQuestion AnsweringSentence

Similar Papers 제목 키워드 기반

ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models

2024-04-17 · Trong-Hieu Nguyen, Anh-Cuong Le, Viet-Cuong Nguyen

The rapid advancement of large language models (LLMs) necessitates the development of new benchmarks to accurately assess their capabilities. To address this need for Vietnamese, this work aims to introduce ViLLM-Eval, t…

Language ModelingLanguage ModellingLarge Language ModelMultiple-choice

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

2026-06-23 · Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat 외 arxiv

Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…

Speaker Recognition

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

2025-11-12 · Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le arxiv

Contemporary Visual Question Answering (VQA) systems remain constrained when confronted with culturally specific content, largely because cultural knowledge is under-represented in training corpora and the reasoning proc…

Visual Question AnsweringObject Detection

A Vietnamese Dataset for Evaluating Machine Reading Comprehension

2020-12-01 · COLING 2020 8 · Kiet Nguyen, Vu Nguyen, Anh Nguyen, Ngan Nguyen

Over 97 million inhabitants speak Vietnamese as the native language in the world. However, there are few research studies on machine reading comprehension (MRC) in Vietnamese, the task of understanding a document or text…

ArticlesMachine Reading ComprehensionQuestion AnsweringReading Comprehension+1

A study of Vietnamese readability assessing through semantic and statistical features

2024-11-07 · Hung Tuan Le, Long Truong To, Manh Trong Nguyen, Quyen Nguyen 외

Determining the difficulty of a text involves assessing various textual features that may impact the reader's text comprehension, yet current research in Vietnamese has only focused on statistical features. This paper in…

Reading Comprehension