paper-with-me

Papers

Textual Data Augmentation for Arabic-English Code-Switching Speech Recognition

2022-01-07 · Amir Hussein, Shammur Absar Chowdhury, Ahmed Abdelali, Najim Dehak, Ahmed Ali, Sanjeev Khudanpur

The pervasiveness of intra-utterance code-switching (CS) in spoken content requires that speech recognition (ASR) systems handle mixed language. Designing a CS-ASR system has many challenges, mainly due to data scarcity, grammatical structure complexity, and domain mismatch. The most common method for addressing CS is to train an ASR system with the available transcribed CS speech, along with monolingual data. In this work, we propose a zero-shot learning methodology for CS-ASR by augmenting the monolingual data with artificially generating CS text. We based our approach on random lexical replacements and Equivalence Constraint (EC) while exploiting aligned translation pairs to generate random and grammatically valid CS content. Our empirical results show a 65.5% relative reduction in language model perplexity, and 7.7% in ASR WER on two ecologically valid CS test sets. The human evaluation of the generated text using EC suggests that more than 80% is of adequate quality.

📄 PDF Abstract BibTeX arXiv:2201.02550

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionText AugmentationTranslationvalidZero-Shot Learning

Similar Papers 제목 키워드 기반

LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect

2025-04-03 · Hedi Naouara, Jean-Pierre Lorré, Jérôme Louradour

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we p…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+2

Computational Approaches to Arabic-English Code-Switching

2024-10-17 · Caroline Sabty

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech…

Data AugmentationLanguage Identificationnamed-entity-recognitionNamed Entity Recognition+4

Accenture at CheckThat! 2021: Interesting claim identification and ranking with contextually sensitive lexical training data augmentation

2021-07-12 · Evan Williams, Paul Rodrigues, Sieu Tran

This paper discusses the approach used by the Accenture Team for CLEF2021 CheckThat! Lab, Task 1, to identify whether a claim made in social media would be interesting to a wide audience and should be fact-checked. Twitt…

Data Augmentation

Data Augmentation for Maltese NLP using Transliterated and Machine Translated Arabic Data

2025-09-16 · Kurt Micallef, Nizar Habash, Claudia Borg arxiv

Maltese is a unique Semitic language that has evolved under extensive influence from Romance and Germanic languages, particularly Italian and English. Despite its Semitic roots, its orthography is based on the Latin scri…

Machine TranslationData Augmentation

A Contextual Word Embedding for Arabic Sarcasm Detection with Random Forests

2021-04-01 · EACL (WANLP) 2021 4 · Hazem Elgabry, Shimaa Attia, Ahmed Abdel-Rahman, Ahmed Abdel-Ate 외

Sarcasm detection is of great importance in understanding people’s true sentiments and opinions. Many online feedbacks, reviews, social media comments, etc. are sarcastic. Several researches have already been done in thi…

Data AugmentationSarcasm Detection