Do not neglect related languages: The case of low-resource Occitan cross-lingual word embeddings
Cross-lingual word embeddings (CLWEs) have proven indispensable for various natural language processing tasks, e.g., bilingual lexicon induction (BLI). However, the lack of data often impairs the quality of representations. Various approaches requiring only weak cross-lingual supervision were proposed, but current methods still fail to learn good CLWEs for languages with only a small monolingual corpus. We therefore claim that it is necessary to explore further datasets to improve CLWEs in low-resource setups. In this paper we propose to incorporate data of related high-resource languages. In contrast to previous approaches which leverage independently pre-trained embeddings of languages, we (i) train CLWEs for the low-resource and a related language jointly and (ii) map them to the target language to build the final multilingual space. In our experiments we focus on Occitan, a low-resource Romance language which is often neglected due to lack of resources. We leverage data from French, Spanish and Catalan for training and evaluate on the Occitan-English BLI task. By incorporating supporting languages our method outperforms previous approaches by a large margin. Furthermore, our analysis shows that the degree of relatedness between an incorporated language and the low-resource language is critically important.
Code (0)
등록된 구현이 없습니다.
Tasks
Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord EmbeddingsSimilar Papers 제목 키워드 기반
TenTrans Multilingual Low-Resource Translation System for WMT21 Indo-European Languages Task
This paper describes TenTrans’ submission to WMT21 Multilingual Low-Resource Translation shared task for the Romance language pairs. This task focuses on improving translation quality from Catalan to Occitan, Romanian an…
Transfer LearningTranslationCUNI systems for WMT21: Multilingual Low-Resource Translation for Indo-European Languages Shared Task
This paper describes Charles University submission for Multilingual Low-Resource Translation for Indo-European Languages shared task at WMT21. We competed in translation from Catalan into Romanian, Italian and Occitan. O…
Grapheme-to-Phoneme ConversionMulti-Task LearningTranslationAutomatic Transcription of Handwritten Old Occitan Language
While existing neural network-based approaches have shown promising results in Handwritten Text Recognition (HTR) for high-resource languages and standardized/machine-written text, their application to low-resource langu…
Data AugmentationDecoderHandwritten Text RecognitionHTRReview on the Existing Language Resources for Languages of France
With the support of the DGLFLF, ELDA conducted an inventory of existing language resources for the regional languages of France. The main aim of this inventory was to assess the exploitability of the identified resources…
Cultural Vocal Bursts Intensity PredictionDiversityTranslationModeling Orthographic Variation in Occitan's Dialects
Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a multilingual model with data from several Oc…
Dependency ParsingPart-Of-Speech Tagging