paper-with-me

홈 › Papers

Do not neglect related languages: The case of low-resource Occitan cross-lingual word embeddings

2021-11-01 · EMNLP (MRL) 2021 11 · Lisa Woller, Viktor Hangya, Alexander Fraser

Cross-lingual word embeddings (CLWEs) have proven indispensable for various natural language processing tasks, e.g., bilingual lexicon induction (BLI). However, the lack of data often impairs the quality of representations. Various approaches requiring only weak cross-lingual supervision were proposed, but current methods still fail to learn good CLWEs for languages with only a small monolingual corpus. We therefore claim that it is necessary to explore further datasets to improve CLWEs in low-resource setups. In this paper we propose to incorporate data of related high-resource languages. In contrast to previous approaches which leverage independently pre-trained embeddings of languages, we (i) train CLWEs for the low-resource and a related language jointly and (ii) map them to the target language to build the final multilingual space. In our experiments we focus on Occitan, a low-resource Romance language which is often neglected due to lack of resources. We leverage data from French, Spanish and Catalan for training and evaluate on the Occitan-English BLI task. By incorporating supporting languages our method outperforms previous approaches by a large margin. Furthermore, our analysis shows that the degree of relatedness between an incorporated language and the low-resource language is critically important.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Bilingual Lexicon InductionCross-Lingual Word EmbeddingsWord Embeddings

Similar Papers 제목 키워드 기반

TenTrans Multilingual Low-Resource Translation System for WMT21 Indo-European Languages Task

2021-11-01 · WMT (EMNLP) 2021 11 · Han Yang, Bojie Hu, Wanying Xie, Ambyera Han 외

This paper describes TenTrans’ submission to WMT21 Multilingual Low-Resource Translation shared task for the Romance language pairs. This task focuses on improving translation quality from Catalan to Occitan, Romanian an…

Transfer LearningTranslation

CUNI systems for WMT21: Multilingual Low-Resource Translation for Indo-European Languages Shared Task

2021-09-20 · WMT (EMNLP) 2021 11 · Josef Jon, Michal Novák, João Paulo Aires, Dušan Variš 외

This paper describes Charles University submission for Multilingual Low-Resource Translation for Indo-European Languages shared task at WMT21. We competed in translation from Catalan into Romanian, Italian and Occitan. O…

Grapheme-to-Phoneme ConversionMulti-Task LearningTranslation

Automatic Transcription of Handwritten Old Occitan Language

2023-12-06 · EMNLP 2023 12 · Esteban Garces Arias, Vallari Pai, Matthias Schöffel, Christian Heumann 외

While existing neural network-based approaches have shown promising results in Handwritten Text Recognition (HTR) for high-resource languages and standardized/machine-written text, their application to low-resource langu…

Data AugmentationDecoderHandwritten Text RecognitionHTR

Review on the Existing Language Resources for Languages of France

2016-05-01 · LREC 2016 5 · Thibault Grouas, Val{\'e}rie Mapelli, Quentin Samier

With the support of the DGLFLF, ELDA conducted an inventory of existing language resources for the regional languages of France. The main aim of this inventory was to assess the exploitability of the identified resources…

Cultural Vocal Bursts Intensity PredictionDiversityTranslation

Modeling Orthographic Variation in Occitan's Dialects

2024-04-30 · Zachary William Hopton, Noëmi Aepli

Effectively normalizing textual data poses a considerable challenge, especially for low-resource languages lacking standardized writing systems. In this study, we fine-tuned a multilingual model with data from several Oc…

Dependency ParsingPart-Of-Speech Tagging