Reading Comprehension in Czech via Machine Translation and Cross-lingual Transfer
Reading comprehension is a well studied task, with huge training datasets in English. This work focuses on building reading comprehension systems for Czech, without requiring any manually annotated Czech training data. First of all, we automatically translated SQuAD 1.1 and SQuAD 2.0 datasets to Czech to create training and development data, which we release at http://hdl.handle.net/11234/1-3249. We then trained and evaluated several BERT and XLM-RoBERTa baseline models. However, our main focus lies in cross-lingual transfer models. We report that a XLM-RoBERTa model trained on English data and evaluated on Czech achieves very competitive performance, only approximately 2 percent points worse than a~model trained on the translated Czech data. This result is extremely good, considering the fact that the model has not seen any Czech data during training. The cross-lingual transfer approach is very flexible and provides a reading comprehension in any language, for which we have enough monolingual raw texts.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual TransferMachine TranslationReading ComprehensionTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Machine Translation from an Intercomprehension Perspective
Within the first shared task on machine translation between similar languages, we present our first attempts on Czech to Polish machine translation from an intercomprehension perspective. We propose methods based on the …
Machine TranslationTranslationCross-lingual Machine Reading Comprehension with Language Branch Knowledge Distillation
Cross-lingual Machine Reading Comprehension (CLMRC) remains a challenging problem due to the lack of large-scale annotated datasets in low-source languages, such as Arabic, Hindi, and Vietnamese. Many previous approaches…
Knowledge DistillationMachine Reading ComprehensionReading ComprehensionTranslationCross-Lingual Machine Reading Comprehension
Though the community has made great progress on Machine Reading Comprehension (MRC) task, most of the previous works are solving English-based MRC problems, and there are few efforts on other languages mainly due to the …
Machine Reading ComprehensionReading ComprehensionTranslationXCMRC: Evaluating Cross-lingual Machine Reading Comprehension
We present XCMRC, the first public cross-lingual language understanding (XLU) benchmark which aims to test machines on their cross-lingual reading comprehension ability. To be specific, XCMRC is a Cross-lingual Cloze-sty…
Machine Reading ComprehensionReading ComprehensionSentenceComprehension of Subtitles from Re-Translating Simultaneous Speech Translation
In simultaneous speech translation, one can vary the size of the output window, system latency and sometimes the allowed level of rewriting. The effect of these properties on readability and comprehensibility has not bee…
Machine TranslationTranslation