paper-with-me

홈 › Papers

MLQA: Evaluating Cross-lingual Extractive Question Answering

2019-10-16 · ACL 2020 6 · Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, Holger Schwenk

Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets. Such annotated datasets are difficult and costly to collect, and rarely exist in languages other than English, making training QA systems in other languages challenging. An alternative to building large monolingual training datasets is to develop cross-lingual systems which can transfer to a target language without requiring training data in that language. In order to develop such systems, it is crucial to invest in high quality multilingual evaluation benchmarks to measure progress. We present MLQA, a multi-way aligned extractive QA evaluation benchmark intended to spur research in this area. MLQA contains QA instances in 7 languages, namely English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese. It consists of over 12K QA instances in English and 5K in each other language, with each QA instance being parallel between 4 languages on average. MLQA is built using a novel alignment context strategy on Wikipedia articles, and serves as a cross-lingual extension to existing extractive QA datasets. We evaluate current state-of-the-art cross-lingual representations on MLQA, and also provide machine-translation-based baselines. In all cases, transfer results are shown to be significantly behind training-language performance.

📄 PDF Abstract BibTeX arXiv:1910.07475

Code (4)

facebookresearch/MLQA 공식 구현
ccasimiro88/TranslateAlignRetrieve
lmarent/TranslateAlignRetrieve
stonybrooknlp/musique

Tasks

ArticlesExtractive Question-AnsweringMachine TranslationQuestion Answering

Similar Papers 제목 키워드 기반

Automatic Spanish Translation of SQuAD Dataset for Multi-lingual Question Answering

2020-05-01 · LREC 2020 5 · Casimiro Pio Carrino, Marta R. Costa-juss{\`a}, Jos{\'e} A. R. Fonollosa

Recently, multilingual question answering became a crucial research topic, and it is receiving increased interest in the NLP community. However, the unavailability of large-scale datasets makes it challenging to train mu…

Question AnsweringTARTranslation

Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering

2019-12-11 · Casimiro Pio Carrino, Marta R. Costa-jussà, José A. R. Fonollosa

Recently, multilingual question answering became a crucial research topic, and it is receiving increased interest in the NLP community. However, the unavailability of large-scale datasets makes it challenging to train mu…

Question AnsweringTARTranslation

Bridging the Language Gap: Knowledge Injected Multilingual Question Answering

2023-04-06 · Zhichao Duan, Xiuxing Li, Zhengyan Zhang, Zhenyu Li 외

Question Answering (QA) is the task of automatically answering questions posed by humans in natural languages. There are different settings to answer a question, such as abstractive, extractive, boolean, and multiple-cho…

Cross-Lingual TransferExtractive Question-AnsweringLink PredictionMultiple-choice+1

Building a Swedish Question-Answering Model

2020-06-01 · PaM 2020 6 · Hannes von Essen, Daniel Hesslow

High quality datasets for question answering exist in a few languages, but far from all. Producing such datasets for new languages requires extensive manual labour. In this work we look at different methods for using exi…

Machine TranslationmodelQuestion Answering

VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation

2025-10-23 · Son T. Luu, Trung Vo, Hiep Nguyen, Khanh Quoc Tran 외 arxiv

This paper presents the VLSP 2025 MLQA-TSR - the multimodal legal question answering on traffic sign regulation shared task at VLSP 2025. VLSP 2025 MLQA-TSR comprises two subtasks: multimodal legal retrieval and multimod…

Question Answering