paper-with-me

홈 › Papers

Linguistic Diversity Scores for NLP Data Sets

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Quantifying linguistic diversity in multilingual data sets is important for improving cross-linguistic coverage of NLP models. However, current linguistic diversity scores rely mostly on measures such as the number of languages in the sample, which are not very informative about the structural properties of languages. In this paper, we propose a score derived from the distribution of text statistics (mean word length) as a linguistic attribute suitable for cross-linguistic comparison. We compare NLP data sets (UD, Bible100. mBERT, XTREME, XGLUE, XNLI, XCOPA, TyDiQA, XQuAD) to a new data set designed specifically for the purpose of being typologically representative (WALS-SC). To do so, we apply a version of the Jaccard index ($J_{mm}$) suitable for comparing sets of measures. This diversity score can identify the types of languages that need to be included in multilingual data sets in order to reach broad linguistic coverage. We find, for example, that (poly)synthetic languages are missing in almost all data sets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDiversity

Methods 이 논문이 사용한 방법론

mBERT mBERT

Similar Papers 제목 키워드 기반

The Ghanaian NLP Landscape: A First Look

2024-05-10 · Sheriff Issaka, Zhaoyi Zhang, Mihir Heda, Keyi Wang 외

Despite comprising one-third of global languages, African languages are critically underrepresented in Artificial Intelligence (AI), threatening linguistic diversity and cultural heritage. Ghanaian languages, in particul…

Diversity

The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text

2023-11-16 · Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, Chloé Clavel

This study investigates the consequences of training language models on synthetic data generated by their predecessors, an increasingly prevalent practice given the prominence of powerful generative models. Diverging fro…

DiversityText Generation

Where Are We? Evaluating LLM Performance on African Languages

2025-02-26 · Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim Elmadany 외

Africa's rich linguistic heritage remains underrepresented in NLP, largely due to historical policies that favor foreign languages and create significant data inequities. In this paper, we integrate theoretical insights …

Diversity

Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity

2025-04-19 · Maria-Teresa De Rosa Palmini, Eva Cetinic

Following the initial excitement, Text-to-Image (TTI) models are now being examined more critically. While much of the discourse has focused on biases and stereotypes embedded in large-scale training datasets, the sociot…

DiversityImage GenerationSemantic SimilaritySemantic Textual Similarity+2

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets

2024-03-06 · Tanja Samardzic, Ximena Gutierrez, Christian Bentz, Steven Moran 외

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families …

DiversityMultilingual NLP