paper-with-me

홈 › Papers

Contextual Embeddings for Arabic-English Code-Switched Data

2020-12-01 · COLING (WANLP) 2020 12 · Caroline Sabty, Mohamed Islam, Slim Abdennadher

Globalization has caused the rise of the code-switching phenomenon among multilingual societies. In Arab countries, code-switching between Arabic and English has become frequent, especially through social media platforms. Consequently, research in Natural Language Processing (NLP) systems increased to tackle such a phenomenon. One of the significant challenges of developing code-switched NLP systems is the lack of data itself. In this paper, we propose an open source trained bilingual contextual word embedding models of FLAIR, BERT, and ELECTRA. We also propose a novel contextual word embedding model called KERMIT, which can efficiently map Arabic and English words inside one vector space in terms of data usage. We applied intrinsic and extrinsic evaluation methods to compare the performance of the models. Our results show that FLAIR and FastText achieve the highest results in the sentiment analysis task. However, KERMIT is the best-achieving model on the intrinsic evaluation and named entity recognition. Also, it outperforms the other transformer-based models on question answering task.

📄 PDF Abstract BibTeX

Code (1)

csabty/code-switch-arabic-english-contextual-embeddings 공식 구현 tf

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Question AnsweringSentiment Analysis

Similar Papers 제목 키워드 기반

Computational Approaches to Arabic-English Code-Switching

2024-10-17 · Caroline Sabty

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech…

Data AugmentationLanguage Identificationnamed-entity-recognitionNamed Entity Recognition+4

ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs

2024-06-26 · Ahmed Heakl, Youssef Zaghloul, Mennatullah Ali, Rania Hossam 외

Motivated by the widespread increase in the phenomenon of code-switching between Egyptian Arabic and English in recent times, this paper explores the intricacies of machine translation (MT) and automatic speech recogniti…

ArzEn Code-switched Translation to araArzEn Code-switched Translation to engArzEn Speech RecognitionAutomatic Speech Recognition+9

ArzEn-ST: A Three-way Speech Translation Corpus for Code-Switched Egyptian Arabic - English

2022-11-22 · Injy Hamed, Nizar Habash, Slim Abdennadher, Ngoc Thang Vu

We present our work on collecting ArzEn-ST, a code-switched Egyptian Arabic - English Speech Translation Corpus. This corpus is an extension of the ArzEn speech corpus, which was collected through informal interviews wit…

Machine TranslationTranslation

Code-Switched Named Entity Recognition with Embedding Attention

2018-07-01 · WS 2018 7 · Changhan Wang, Kyunghyun Cho, Douwe Kiela

We describe our work for the CALCS 2018 shared task on named entity recognition on code-switched data. Our system ranked first place for MS Arabic-Egyptian named entity recognition and third place for English-Spanish.

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

Tackling Code-Switched NER: Participation of CMU

2018-07-01 · WS 2018 7 · Parvathy Geetha, Ch, Khyathi u, Alan W. black

Named Entity Recognition plays a major role in several downstream applications in NLP. Though this task has been heavily studied in formal monolingual texts and also noisy texts like Twitter data, it is still an emerging…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1