paper-with-me

홈 › Papers

Are pre-trained text representations useful for multilingual and multi-dimensional language proficiency modeling?

2021-02-25 · Taraka Rama, Sowmya Vajjala

Development of language proficiency models for non-native learners has been an active area of interest in NLP research for the past few years. Although language proficiency is multidimensional in nature, existing research typically considers a single "overall proficiency" while building models. Further, existing approaches also considers only one language at a time. This paper describes our experiments and observations about the role of pre-trained and fine-tuned multilingual embeddings in performing multi-dimensional, multilingual language proficiency classification. We report experiments with three languages -- German, Italian, and Czech -- and model seven dimensions of proficiency ranging from vocabulary control to sociolinguistic appropriateness. Our results indicate that while fine-tuned embeddings are useful for multilingual proficiency modeling, none of the features achieve consistently best performance for all dimensions of language proficiency. All code, data and related supplementary material can be found at: https://github.com/nishkalavallabhi/MultidimCEFRScoring.

📄 PDF Abstract BibTeX arXiv:2102.12971

Code (1)

nishkalavallabhi/MultidimCEFRScoring 공식 구현 pytorch

Similar Papers 제목 키워드 기반

On the Language Neutrality of Pre-trained Multilingual Representations

2020-04-09 · Findings of the Association for Computational Linguistics 2020 · Jindřich Libovický, Rudolf Rosa, Alexander Fraser

Multilingual contextual embeddings, such as multilingual BERT and XLM-RoBERTa, have proved useful for many multi-lingual tasks. Previous work probed the cross-linguality of the representations indirectly using zero-shot …

Language IdentificationTransfer LearningWord Alignment

Multilingual Alignment of Contextual Word Representations

2020-02-10 · ICLR 2020 1 · Steven Cao, Nikita Kitaev, Dan Klein

We propose procedures for evaluating and strengthening contextual embedding alignment and show that they are useful in analyzing and improving multilingual BERT. In particular, after our proposed alignment procedure, BER…

Retrieval

Distilling a Pretrained Language Model to a Multilingual ASR Model

2022-06-25 · Kwanghee Choi, Hyung-Min Park

Multilingual speech data often suffer from long-tailed language distribution, resulting in performance degradation. However, multilingual text data is much easier to obtain, yielding a more useful general language model.…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3

Cross-lingual intent classification in a low resource industrial setting

2019-11-01 · IJCNLP 2019 11 · Talaat Khalil, Kornel Kie{\l}czewski, Georgios Christos Chouliaras, Amina Keldibek 외

This paper explores different approaches to multilingual intent classification in a low resource setting. Recent advances in multilingual text representations promise cross-lingual transfer for classifiers. We investigat…

ClassificationCross-Lingual TransferGeneral Classificationintent-classification+1

Assessing Multilingual Fairness in Pre-trained Multimodal Representations

2021-06-12 · Findings (ACL) 2022 5 · Jialu Wang, Yang Liu, Xin Eric Wang

Recently pre-trained multimodal models, such as CLIP, have shown exceptional capabilities towards connecting images and natural language. The textual representations in English can be desirably transferred to multilingua…

Fairness