Linguistic Diversity Scores for NLP Data Sets
Quantifying linguistic diversity in multilingual data sets is important for improving cross-linguistic coverage of NLP models. However, current linguistic diversity scores rely mostly on measures such as the number of languages in the sample, which are not very informative about the structural properties of languages. In this paper, we propose a score derived from the distribution of text statistics (mean word length) as a linguistic attribute suitable for cross-linguistic comparison. We compare NLP data sets (UD, Bible100. mBERT, XTREME, XGLUE, XNLI, XCOPA, TyDiQA, XQuAD) to a new data set designed specifically for the purpose of being typologically representative (WALS-SC). To do so, we apply a version of the Jaccard index ($J_{mm}$) suitable for comparing sets of measures. This diversity score can identify the types of languages that need to be included in multilingual data sets in order to reach broad linguistic coverage. We find, for example, that (poly)synthetic languages are missing in almost all data sets.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeDiversityMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The Ghanaian NLP Landscape: A First Look
Despite comprising one-third of global languages, African languages are critically underrepresented in Artificial Intelligence (AI), threatening linguistic diversity and cultural heritage. Ghanaian languages, in particul…
DiversityThe Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text
This study investigates the consequences of training language models on synthetic data generated by their predecessors, an increasingly prevalent practice given the prominence of powerful generative models. Diverging fro…
DiversityText GenerationWhere Are We? Evaluating LLM Performance on African Languages
Africa's rich linguistic heritage remains underrepresented in NLP, largely due to historical policies that favor foreign languages and create significant data inequities. In this paper, we integrate theoretical insights …
DiversityExploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity
Following the initial excitement, Text-to-Image (TTI) models are now being examined more critically. While much of the discourse has focused on biases and stereotypes embedded in large-scale training datasets, the sociot…
DiversityImage GenerationSemantic SimilaritySemantic Textual Similarity+2A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets
Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families …
DiversityMultilingual NLP