paper-with-me

Papers

Natural Language Processing for Dialects of a Language: A Survey

2024-01-11 · Aditya Joshi, Raj Dabre, Diptesh Kanojia, Zhuang Li, Haolan Zhan, Gholamreza Haffari, Doris Dippold

State-of-the-art natural language processing (NLP) models are trained on massive training corpora, and report a superlative performance on evaluation datasets. This survey delves into an important attribute of these datasets: the dialect of a language. Motivated by the performance degradation of NLP models for dialectal datasets and its implications for the equity of language technologies, we survey past research in NLP for dialects in terms of datasets, and approaches. We describe a wide range of NLP tasks in terms of two categories: natural language understanding (NLU) (for tasks such as dialect classification, sentiment analysis, parsing, and NLU benchmarks) and natural language generation (NLG) (for summarisation, machine translation, and dialogue systems). The survey is also broad in its coverage of languages which include English, Arabic, German, among others. We observe that past work in NLP concerning dialects goes deeper than mere dialect classification, and extends to several NLU and NLG tasks. For these tasks, we describe classical machine learning using statistical models, along with the recent deep learning-based approaches based on pre-trained language models. We expect that this survey will be useful to NLP researchers interested in building equitable language technologies by rethinking LLM benchmarks and model architectures.

📄 PDF Abstract BibTeX arXiv:2401.05632

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeMachine TranslationNatural Language UnderstandingSentenceSentiment AnalysisSurveyText Generation

Similar Papers 제목 키워드 기반

What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects

2024-02-19 · Verena Blaschke, Christoph Purschke, Hinrich Schütze, Barbara Plank

Natural language processing (NLP) has largely focused on modelling standardized languages. More recently, attention has increasingly shifted to local, non-standardized languages and dialects. However, the relevant speake…

Machine Translation

NLP for The Greek Language: A Longer Survey

2024-08-20 · Katerina Papantoniou, Yannis Tzitzikas

English language is in the spotlight of the Natural Language Processing (NLP) community with other languages, like Greek, lagging behind in terms of offered methods, tools and resources. Due to the increasing interest in…

Information RetrievalManagementRetrievalSurvey

Building a Corpus of Qatari Arabic Expressions

2020-05-01 · LREC 2020 5 · Sara Al-Mulla, Wajdi Zaghouani

The current Arabic natural language processing resources are mainly build to address the Modern Standard Arabic (MSA), while we witnessed some scattered efforts to build resources for various Arabic dialects such as the …

Automatic Arabic Dialect Identification Systems for Written Texts: A Survey

2020-09-26 · Maha J. Althobaiti

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…

Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6

A Survey of Corpora for Germanic Low-Resource Languages and Dialects

2023-04-19 · Verena Blaschke, Hinrich Schütze, Barbara Plank

Despite much progress in recent years, the vast majority of work in natural language processing (NLP) is on standard languages with many speakers. In this work, we instead focus on low-resource languages and in particula…

Survey