paper-with-me

홈 › Papers

A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets

2024-03-06 · Tanja Samardzic, Ximena Gutierrez, Christian Bentz, Steven Moran, Olga Pelloni

Typologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP. Linguistic diversity of these data sets is typically measured as the number of languages or language families included in the sample, but such measures do not consider structural properties of the included languages. In this paper, we propose assessing linguistic diversity of a data set against a reference language sample as a means of maximising linguistic diversity in the long run. We represent languages as sets of features and apply a version of the Jaccard index suitable for comparing sets of measures. In addition to the features extracted from typological data bases, we propose an automatic text-based measure, which can be used as a means of overcoming the well-known problem of data sparsity in manually collected features. Our diversity score is interpretable in terms of linguistic features and can identify the types of languages that are not represented in a data set. Using our method, we analyse a range of popular multilingual data sets (UD, Bible100, mBERT, XTREME, XGLUE, XNLI, XCOPA, TyDiQA, XQuAD). In addition to ranking these data sets, we find, for example, that (poly)synthetic languages are missing in almost all of them.

📄 PDF Abstract BibTeX arXiv:2403.03909

Code (1)

morphdiv/jmm_diversity 공식 구현

Tasks

DiversityMultilingual NLP

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
mBERT mBERT

Similar Papers 제목 키워드 기반

Linguistic Diversity Scores for NLP Data Sets

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Quantifying linguistic diversity in multilingual data sets is important for improving cross-linguistic coverage of NLP models. However, current linguistic diversity scores rely mostly on measures such as the number of la…

AttributeDiversity

TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP

2022-06-01 · LREC 2022 6 · Steven Moran, Christian Bentz, Ximena Gutierrez-Vasques, Olga Pelloni 외

We present the TeDDi sample, a diversity sample of text data for language comparison and multilingual Natural Language Processing. The TeDDi sample currently features 89 languages based on the typological diversity sampl…

DiversityMultilingual NLP

CTAP for Italian: Integrating Components for the Analysis of Italian into a Multilingual Linguistic Complexity Analysis Tool

2020-05-01 · LREC 2020 5 · Nadezda Okinina, Jennifer-Carmen Frey, Zarah Weiss

Linguistic complexity research being a very actively developing field, an increasing number of text analysis tools are created that use natural language processing techniques for the automatic extraction of quantifiable …

Diversity

X-Guard: Multilingual Guard Agent for Content Moderation

2025-04-11 · Bibek Upadhayay, Vahid Behzadan, Ph. D

Large Language Models (LLMs) have rapidly become integral to numerous applications in critical domains where reliability is paramount. Despite significant advances in safety frameworks and guardrails, current protective …

Differential Privacy, Linguistic Fairness, and Training Data Influence: Impossibility and Possibility Theorems for Multilingual Language Models

2023-08-17 · Phillip Rust, Anders Søgaard

Language models such as mBERT, XLM-R, and BLOOM aim to achieve multilingual generalization or compression to facilitate transfer to a large number of (potentially unseen) languages. However, these models should ideally a…

FairnessXLM-R