paper-with-me

홈 › Papers

Investigating Code-Mixed Modern Standard Arabic-Egyptian to English Machine Translation

2021-05-28 · NAACL (CALCS) 2021 6 · El Moatez Billah Nagoudi, AbdelRahim Elmadany, Muhammad Abdul-Mageed

Recent progress in neural machine translation (NMT) has made it possible to translate successfully between monolingual language pairs where large parallel data exist, with pre-trained models improving performance even further. Although there exists work on translating in code-mixed settings (where one of the pairs includes text from two or more languages), it is still unclear what recent success in NMT and language modeling exactly means for translating code-mixed text. We investigate one such context, namely MT from code-mixed Modern Standard Arabic and Egyptian Arabic (MSAEA) into English. We develop models under different conditions, employing both (i) standard end-to-end sequence-to-sequence (S2S) Transformers trained from scratch and (ii) pre-trained S2S language models (LMs). We are able to acquire reasonable performance using only MSA-EN parallel data with S2S models trained from scratch. We also find LMs fine-tuned on data from various Arabic dialects to help the MSAEA-EN task. Our work is in the context of the Shared Task on Machine Translation in Code-Switching. Our best model achieves $\bf25.72$ BLEU, placing us first on the official shared task evaluation for MSAEA-EN.

📄 PDF Abstract BibTeX arXiv:2105.13573

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Arabizi Detection and Conversion to Arabic

2013-06-28 · WS 2014 10 · Kareem Darwish

Arabizi is Arabic text that is written using Latin characters. Arabizi is used to present both Modern Standard Arabic (MSA) or Arabic dialects. It is commonly used in informal settings such as social networking sites and…

Language ModelingLanguage ModellingTransliteration

A Conventional Orthography for Tunisian Arabic

2014-05-01 · LREC 2014 5 · In{\`e}s Zribi, Rahma Boujelbane, Abir Masmoudi, Mariem ellouze 외

Tunisian Arabic is a dialect of the Arabic language spoken in Tunisia. Tunisian Arabic is an under-resourced language. It has neither a standard orthography nor large collections of written text and dictionaries. Actuall…

Language ModellingMachine TranslationSpeech RecognitionSpeech Synthesis+1

Computational Approaches to Arabic-English Code-Switching

2024-10-17 · Caroline Sabty

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech…

Data AugmentationLanguage Identificationnamed-entity-recognitionNamed Entity Recognition+4

BERT-based Multi-Task Model for Country and Province Level Modern Standard Arabic and Dialectal Arabic Identification

2021-06-23 · Abdellah El Mekki, Abdelkader El Mahdaouy, Kabil Essefar, Nabil El Mamoun 외

Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task …

Language IdentificationMulti-Task Learning

Simplified guidelines for the creation of Large Scale Dialectal Arabic Annotations

2012-05-01 · LREC 2012 5 · Heba Elfardy, Mona Diab

The Arabic language is a collection of dialectal variants along with the standard form, Modern Standard Arabic (MSA). MSA is used in official Settings while the dialectal variants (DA) correspond to the native tongue of …

Speech Recognition