paper-with-me

Papers

ZAEBUC-Spoken: A Multilingual Multidialectal Arabic-English Speech Corpus

2024-03-27 · Injy Hamed, Fadhl Eryani, David Palfreyman, Nizar Habash

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain topic and then discuss it with an Interlocutor. The meetings cover different topics and are divided into phases with different language setups. The corpus presents a challenging set for automatic speech recognition (ASR), including two languages (Arabic and English) with Arabic spoken in multiple variants (Modern Standard Arabic, Gulf Arabic, and Egyptian Arabic) and English used with various accents. Adding to the complexity of the corpus, there is also code-switching between these languages and dialects. As part of our work, we take inspiration from established sets of transcription guidelines to present a set of guidelines handling issues of conversational speech, code-switching and orthography of both languages. We further enrich the corpus with two layers of annotations; (1) dialectness level annotation for the portion of the corpus where mixing occurs between different variants of Arabic, and (2) automatic morphological annotations, including tokenization, lemmatization, and part-of-speech tagging.

📄 PDF Abstract BibTeX arXiv:2403.18182

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)LemmatizationPart-Of-Speech Taggingspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

A Multidialectal Parallel Corpus of Arabic

2014-05-01 · LREC 2014 5 · Houda Bouamor, Nizar Habash, Kemal Oflazer

The daily spoken variety of Arabic is often termed the colloquial or dialect form of Arabic. There are many Arabic dialects across the Arab World and within other Arabic speaking communities. These dialects vary widely f…

Dialect IdentificationMachine TranslationTranslation

ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus

2022-06-01 · LREC 2022 6 · Nizar Habash, David Palfreyman

We present ZAEBUC, an annotated Arabic-English bilingual writer corpus comprising short essays by first-year university students at Zayed University in the United Arab Emirates. We describe and discuss the various guidel…

LemmatizationPart-Of-Speech TaggingPOSPOS Tagging

AraSAS: The Open Source Arabic Semantic Tagger

2022-06-01 · OSACT (LREC) 2022 6 · Mahmoud El-Haj, Elvis de Souza, Nouran Khallaf, Paul Rayson 외

This paper presents (AraSAS) the first open-source Arabic semantic analysis tagging system. AraSAS is a software framework that provides full semantic tagging of text written in Arabic. AraSAS is based on the UCREL Seman…

DiversityTAG

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task

2025-09-02 · Bashar Talafha, Hawau Olamide Toyin, Peter Sullivan, AbdelRahim Elmadany 외 arxiv

We present the findings of the sixth Nuanced Arabic Dialect Identification (NADI 2025) Shared Task, which focused on Arabic speech dialect processing across three subtasks: spoken dialect identification (Subtask 1), spee…

Speech Recognition

Aladdin-FTI @ AMIYA Three Wishes for Arabic NLP: Fidelity, Diglossia, and Multidialectal Generation

2026-02-18 · Jonathan Mutal, Perla Al Almaoui, Simon Hengchen, Pierrette Bouillon arxiv

Arabic dialects have long been under-represented in Natural Language Processing (NLP) research due to their non-standardization and high variability, which pose challenges for computational modeling. Recent advances in t…

Text Generation