EENLP: Cross-lingual Eastern European NLP Index
Motivated by the sparsity of NLP resources for Eastern European languages, we present a broad index of existing Eastern European language resources (90+ datasets and 45+ models) published as a github repository open for updates from the community. Furthermore, to support the evaluation of commonsense reasoning tasks, we provide hand-crafted cross-lingual datasets for five different semantic tasks (namely news categorization, paraphrase detection, Natural Language Inference (NLI) task, tweet sentiment detection, and news sentiment detection) for some of the Eastern European languages. We perform several experiments with the existing multilingual models on these datasets to define the performance baselines and compare them to the existing results for other languages.
Code (1)
Tasks
Cross-Lingual TransferNatural Language InferenceTransfer LearningSimilar Papers 제목 키워드 기반
Is a Prestigious Job the same as a Prestigious Country? A Case Study on Multilingual Sentence Embeddings and European Countries
We study how multilingual sentence representations capture European countries and occupations and how this differs across European languages. We prompt the models with templated sentences that we machine-translate into 1…
SentenceSentence EmbeddingsIntegration into \'economie-monde and regionalisation of the Central Eastern European space since 1989
The fall of the Berlin Wall in 1989, modified the relations between cities of the former communist bloc. The European and worldwide reorientation of interactions that followed raises the question of the actual state of h…
MultiSlav: Using Cross-Lingual Knowledge Transfer to Combat the Curse of Multilinguality
Does multilingual Neural Machine Translation (NMT) lead to The Curse of the Multlinguality or provides the Cross-lingual Knowledge Transfer within a language family? In this study, we explore multiple approaches for exte…
Machine TranslationNMTTransfer LearningTranslationMassive migration from the steppe is a source for Indo-European languages in Europe
We generated genome-wide data from 69 Europeans who lived between 8,000-3,000 years ago by enriching ancient DNA libraries for a target set of almost four hundred thousand polymorphisms. Enrichment of these positions dec…
DNA analysisFindings of the WMT 2021 Shared Task on Large-Scale Multilingual Machine Translation
We present the results of the first task on Large-Scale Multilingual Machine Translation. The task consists on the many-to-many evaluation of a single model across a variety of source and target languages. This year, the…
Machine TranslationTranslation