paper-with-me

Papers

DIALECTBENCH: A NLP Benchmark for Dialects, Varieties, and Closely-Related Languages

2024-03-16 · Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, Antonios Anastasopoulos

Language technologies should be judged on their usefulness in real-world use cases. An often overlooked aspect in natural language processing (NLP) research and evaluation is language variation in the form of non-standard dialects or language varieties (hereafter, varieties). Most NLP benchmarks are limited to standard language varieties. To fill this gap, we propose DIALECTBENCH, the first-ever large-scale benchmark for NLP on varieties, which aggregates an extensive set of task-varied variety datasets (10 text-level tasks covering 281 varieties). This allows for a comprehensive evaluation of NLP system performance on different language varieties. We provide substantial evidence of performance disparities between standard and non-standard language varieties, and we also identify language clusters with large performance divergence across tasks. We believe DIALECTBENCH provides a comprehensive view of the current state of NLP for language varieties and one step towards advancing it further. Code/data: https://github.com/ffaisal93/DialectBench

📄 PDF Abstract BibTeX arXiv:2403.11009

Code (1)

ffaisal93/dialectbench 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Proceedings of the Joint Workshop on Language Technology for Closely Related Languages, Varieties and Dialects

2015-09-01 · WS 2015 9 ·

Improving Zero-shot Cross-lingual Transfer between Closely Related Languages by injecting Character-level Noise

2021-09-14 · Findings (ACL) 2022 5 · Noëmi Aepli, Rico Sennrich

Cross-lingual transfer between a high-resource language and its dialects or closely related language varieties should be facilitated by their similarity. However, current approaches that operate in the embedding space do…

Cross-Lingual TransferPOSPOS TaggingZero-Shot Cross-Lingual Transfer

When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models

2025-08-18 · Ahmed Elshabrawy, Hour Kaing, Haiyue Song, Alham Fikri Aji 외 arxiv

Alignment with high-resource standard languages is often assumed to aid the modeling of related low-resource varieties. We challenge this assumption by demonstrating that excessive representational entanglement with a do…

DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models

2025-01-27 · Niyati Bafna, Emily Chang, Nathaniel R. Robinson, David R. Mortensen 외

Most of the world's languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ …

Machine Translation

Machine Translation into Low-resource Language Varieties

2021-06-12 · ACL 2021 5 · Sachin Kumar, Antonios Anastasopoulos, Shuly Wintner, Yulia Tsvetkov

State-of-the-art machine translation (MT) systems are typically trained to generate the "standard" target language; however, many languages have multiple varieties (regional varieties, dialects, sociolects, non-native va…

Machine TranslationTranslation