paper-with-me

홈 › Papers

What is "Typological Diversity" in NLP?

2024-02-06 · Esther Ploeger, Wessel Poelman, Miryam de Lhoneux, Johannes Bjerva

The NLP research community has devoted increased attention to languages beyond English, resulting in considerable improvements for multilingual NLP. However, these improvements only apply to a small subset of the world's languages. Aiming to extend this, an increasing number of papers aspires to enhance generalizable multilingual performance across languages. To this end, linguistic typology is commonly used to motivate language selection, on the basis that a broad typological sample ought to imply generalization across a broad range of languages. These selections are often described as being 'typologically diverse'. In this work, we systematically investigate NLP research that includes claims regarding 'typological diversity'. We find there are no set definitions or criteria for such claims. We introduce metrics to approximate the diversity of language selection along several axes and find that the results vary considerably across papers. Crucially, we show that skewed language selection can lead to overestimated multilingual performance. We recommend future work to include an operationalization of 'typological diversity' that empirically justifies the diversity of language samples.

📄 PDF Abstract BibTeX arXiv:2402.04222

Code (2)

wpoelman/typ-div 공식 구현
wpoelman/typ-div-survey 공식 구현

Tasks

DiversityMultilingual NLP

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Overlooked Data in Typological Databases: What Grambank Teaches Us About Gaps in Grammars

2022-06-01 · LREC 2022 6 · Jakob Lesage, Hannah J. Haynie, Hedvig Skirgård, Tobias Weber 외

Typological databases can contain a wealth of information beyond the collection of linguistic properties across languages. This paper shows how information often overlooked in typological databases can inform the researc…

DescriptiveDiversityNegation

Dialects of Translationese Shape Language Model Learning

2026-02-18 · Jenny Kunz arxiv

Machine-translated data is widely used in multilingual NLP, particularly where native text is scarce. However, translated text differs systematically from native text. This phenomenon is known as translationese, and it r…

Linguistic AcceptabilityLanguage Modelling

Towards Afrocentric NLP for African Languages: Where We Are and Where We Can Go

2022-03-16 · ACL 2022 5 · Ife Adebara, Muhammad Abdul-Mageed

Aligning with ACL 2022 special Theme on "Language Diversity: from Low Resource to Endangered Languages", we discuss the major linguistic and sociopolitical challenges facing development of NLP technologies for African la…

Diversity

Does Typological Blinding Impede Cross-Lingual Sharing?

2021-01-28 · EACL 2021 2 · Johannes Bjerva, Isabelle Augenstein

Bridging the performance gap between high- and low-resource languages has been the focus of much previous work. Typological features from databases such as the World Atlas of Language Structures (WALS) are a prime candid…

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets

2024-03-06 · Tanja Samardzic, Ximena Gutierrez, Christian Bentz, Steven Moran 외

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families …

DiversityMultilingual NLP