POS Tagging for Improving Code-Switching Identification in Arabic
When speakers code-switch between their native language and a second language or language variant, they follow a syntactic pattern where words and phrases from the embedded language are inserted into the matrix language. This paper explores the possibility of utilizing this pattern in improving code-switching identification between Modern Standard Arabic (MSA) and Egyptian Arabic (EA). We try to answer the question of how strong is the POS signal in word-level code-switching identification. We build a deep learning model enriched with linguistic features (including POS tags) that outperforms the state-of-the-art results by 1.9{\%} on the development set and 1.0{\%} on the test set. We also show that in intra-sentential code-switching, the selection of lexical items is constrained by POS categories, where function words tend to come more often from the dialectal language while the majority of content words come from the standard language.
Code (0)
등록된 구현이 없습니다.
Tasks
POSPOS TaggingSimilar Papers 제목 키워드 기반
Leveraging Pretrained Word Embeddings for Part-of-Speech Tagging of Code Switching Data
Linguistic Code Switching (CS) is a phenomenon that occurs when multilingual speakers alternate between two or more languages/dialects within a single conversation. Processing CS data is especially challenging in intra-s…
Part-Of-Speech TaggingPOSPOS TaggingSingle Particle Analysis+1LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation
Recent trends in NLP research have raised an interest in linguistic code-switching (CS); modern approaches have been proposed to solve a wide range of NLP tasks on multiple language pairs. Unfortunately, these proposed m…
Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2Arabic Dialect Identification in the Context of Bivalency and Code-Switching
ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus
We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LemmatizationPart-Of-Speech Tagging+2A Tale of Two Scripts: Transliteration and Post-Correction for Judeo-Arabic
Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for J…
Machine Translation