Cross-lingual Transfer Learning with Data Selection for Large-Scale Spoken Language Understanding
A typical cross-lingual transfer learning approach boosting model performance on a language is to pre-train the model on all available supervised data from another language. However, in large-scale systems this leads to high training times and computational requirements. In addition, characteristic differences between the source and target languages raise a natural question of whether source data selection can improve the knowledge transfer. In this paper, we address this question and propose a simple but effective language model based source-language data selection method for cross-lingual transfer learning in large-scale spoken language understanding. The experimental results show that with data selection i) source data and hence training speed is reduced significantly and ii) model performance is improved.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual TransferLanguage ModelingLanguage ModellingSpoken Language UnderstandingTransfer LearningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages
Cross-lingual transfer learning with large multilingual pre-trained models can be an effective approach for low-resource languages with no labeled training data. Existing evaluations of cross-lingual generalisability of …
Cross-Lingual TransferPart-Of-Speech TaggingPOSPOS Tagging+1Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages
Cross-lingual transfer learning with large multilingual pre-trained models can be an effective approach for low-resource languages with no labeled training data. Existing evaluations of zero-shot cross-lingual generalisa…
Cross-Lingual TransferPart-Of-Speech TaggingPOSPOS Tagging+2Budget-Xfer: Budget-Constrained Source Language Selection for Cross-Lingual Transfer to African Languages
Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data from higher-resource sources, yet existing comparisons of source language selection strategies do not control for total tr…
Cross-Lingual TransferSentiment AnalysisNLPDove at SemEval-2020 Task 12: Improving Offensive Language Detection with Cross-lingual Transfer
This paper describes our approach to the task of identifying offensive languages in a multilingual setting. We investigate two data augmentation strategies: using additional semi-supervised labels with different threshol…
Cross-Lingual TransferData AugmentationLanguage IdentificationTranslationGeneralised Unsupervised Domain Adaptation of Neural Machine Translation with Cross-Lingual Data Selection
This paper considers the unsupervised domain adaptation problem for neural machine translation (NMT), where we assume the access to only monolingual text in either the source or target language in the new domain. We prop…
Contrastive LearningDomain AdaptationMachine TranslationNMT+2