paper-with-me

Papers

Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi

2020-05-01 · Benjamin Muller, Benoit Sagot, Djamé Seddah

Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained language models provides new modeling tools to tackle this. In this work, we study the ability of multilingual language models to process an unseen dialect. We take user generated North-African Arabic as our case study, a resource-poor dialectal variety of Arabic with frequent code-mixing with French and written in Arabizi, a non-standardized transliteration of Arabic to Latin script. Focusing on two tasks, part-of-speech tagging and dependency parsing, we show in zero-shot and unsupervised adaptation scenarios that multilingual language models are able to transfer to such an unseen dialect, specifically in two extreme cases: (i) across scripts, using Modern Standard Arabic as a source language, and (ii) from a distantly related language, unseen during pretraining, namely Maltese. Our results constitute the first successful transfer experiments on this dialect, paving thus the way for the development of an NLP ecosystem for resource-scarce, non-standardized and highly variable vernacular languages.

📄 PDF Abstract BibTeX arXiv:2005.00318

Code (0)

등록된 구현이 없습니다.

Tasks

Dependency ParsingPart-Of-Speech TaggingTransliteration

Similar Papers 제목 키워드 기반

Using natural language prompts for machine translation

2022-02-23 · Xavier Garcia, Orhan Firat

We explore the use of natural language prompts for controlling various aspects of the outputs generated by machine translation models. We demonstrate that natural language prompts allow us to influence properties like fo…

Machine TranslationTranslation

Fine-Tuning BERT with Character-Level Noise for Zero-Shot Transfer to Dialects and Closely-Related Languages

2023-03-30 · Aarohi Srivastava, David Chiang

In this work, we induce character-level noise in various forms when fine-tuning BERT to enable zero-shot cross-lingual transfer to unseen dialects and languages. We fine-tune BERT on three sentence-level classification t…

Cross-Lingual TransferSentenceZero-Shot Cross-Lingual Transfer

Incorporating Dialectal Variability for Socially Equitable Language Identification

2017-07-01 · ACL 2017 7 · David Jurgens, Yulia Tsvetkov, Dan Jurafsky

Language identification (LID) is a critical first step for processing multilingual text. Yet most LID systems are not designed to handle the linguistic diversity of global platforms like Twitter, where local dialects and…

DiversityLanguage Identification

SIB-200: A Simple, Inclusive, and Big Evaluation Dataset for Topic Classification in 200+ Languages and Dialects

2023-09-14 · David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev 외

Despite the progress we have recorded in the last few years in multilingual natural language processing, evaluation is typically limited to a small set of languages with available datasets which excludes a large number o…

Cross-Lingual TransferLanguage ModellingLarge Language ModelMachine Translation+3

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

2023-06-05 · Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed

Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whis…

Arabic Speech RecognitionBenchmarkingspeech-recognitionSpeech Recognition