Transfer to a Low-Resource Language via Close Relatives: The Case Study on Faroese
Multilingual language models have pushed state-of-the-art in cross-lingual NLP transfer. The majority of zero-shot cross-lingual transfer, however, use one and the same massively multilingual transformer (e.g., mBERT or XLM-R) to transfer to all target languages, irrespective of their typological, etymological, and phylogenetic relations to other languages. In particular, readily available data and models of resource-rich sibling languages are often ignored. In this work, we empirically show, in a case study for Faroese -- a low-resource language from a high-resource language family -- that by leveraging the phylogenetic information and departing from the 'one-size-fits-all' paradigm, one can improve cross-lingual transfer to low-resource languages. In particular, we leverage abundant resources of other Scandinavian languages (i.e., Danish, Norwegian, Swedish, and Icelandic) for the benefit of Faroese. Our evaluation results show that we can substantially improve the transfer performance to Faroese by exploiting data and models of closely-related high-resource languages. Further, we release a new web corpus of Faroese and Faroese datasets for named entity recognition (NER), semantic text similarity (STS), and new language models trained on all Scandinavian languages.
Code (0)
등록된 구현이 없습니다.
Tasks
Cross-Lingual Transfernamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NERSTStext similarityXLM-RZero-Shot Cross-Lingual TransferMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Delexicalized Cross-lingual Dependency Parsing for Xibe
Manually annotating a treebank is time-consuming and labor-intensive. We conduct delexicalized cross-lingual dependency parsing experiments, where we train the parser on one language and test on our target language. As o…
Dependency ParsingData Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data
Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin scri…
Machine TranslationData AugmentationEvaluating zero-shot transfers and multilingual models for dependency parsing and POS tagging within the low-resource language family Tupían
This work presents two experiments with the goal of replicating the transferability of dependency parsers and POS taggers trained on closely related languages within the low-resource language family Tupían. The experimen…
Dependency ParsingPOSPOS TaggingTransfer LearningThe bat coronavirus RmYN02 is characterized by a 6-nucleotide deletion at the S1/S2 junction, and its claimed PAA insertion is highly doubtful
Zhou et al. reported the discovery of RmYN02, a strain closely related to SARS-CoV-2, which is claimed to contain a natural PAA amino acid insertion at the S1/S2 junction of the spike protein at the same position of the …
How human-derived brain organoids are built differently from brain organoids derived from genetically-close relatives: A multi-scale hypothesis
How genes affect tissue scale organization remains a longstanding biological puzzle. As experimental efforts aim to quantify gene expression, chromatin organization, cellular structure, and tissue structure, computationa…