paper-with-me

홈 › Papers

The Impact of Code-switched Synthetic Data Quality is Task Dependent: Insights from MT and ASR

2025-03-30 · Injy Hamed, Ngoc Thang Vu, Nizar Habash

Code-switching, the act of alternating between languages, emerged as a prevalent global phenomenon that needs to be addressed for building user-friendly language technologies. A main bottleneck in this pursuit is data scarcity, motivating research in the direction of code-switched data augmentation. However, current literature lacks comprehensive studies that enable us to understand the relation between the quality of synthetic data and improvements on NLP tasks. We extend previous research conducted in this direction on machine translation (MT) with results on automatic speech recognition (ASR) and cascaded speech translation (ST) to test generalizability of findings. Our experiments involve a wide range of augmentation techniques, covering lexical replacements, linguistic theories, and back-translation. Based on the results of MT, ASR, and ST, we draw conclusions and insights regarding the efficacy of various augmentation techniques and the impact of quality on performance.

📄 PDF Abstract BibTeX arXiv:2503.23576

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationMachine Translationspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text

2021-07-14 · ACL 2021 5 · Ishan Tarunesh, Syamantak Kumar, Preethi Jyothi

Generating code-switched text is a problem of growing interest, especially given the scarcity of corpora containing large volumes of real code-switched text. In this work, we adapt a state-of-the-art neural machine trans…

Data AugmentationLanguage ModelingLanguage ModellingMachine Translation+2

The Effect of Alignment Objectives on Code-Switching Translation

2023-09-10 · Mohamed Anwar

One of the things that need to change when it comes to machine translation is the models' ability to translate code-switching content, especially with the rise of social media and user-generated content. In this paper, w…

Machine TranslationTranslation

Breaking Language Barriers: Equitable Performance in Multilingual Language Models

2025-08-18 · Tanay Nagar, Grigorii Khvatskii, Anna Sokol, Nitesh V. Chawla arxiv

Cutting-edge LLMs have emerged as powerful tools for multilingual communication and understanding. However, LLMs perform worse in Common Sense Reasoning (CSR) tasks when prompted in low-resource languages (LRLs) like Hin…

Common Sense Reasoning

Improved Sentiment Detection via Label Transfer from Monolingual to Synthetic Code-Switched Text

2019-06-13 · ACL 2019 7 · Bidisha Samanta, Niloy Ganguly, Soumen Chakrabarti

Multilingual writers and speakers often alternate between two languages in a single discourse, a practice called "code-switching". Existing sentiment detection methods are usually trained on sentiment-labeled monolingual…

Hate Speech Detection

Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?

2026-01-12 · Genta Indra Winata, David Anugraha, Patrick Amadeus Irawan, Anirban Das 외 arxiv

Code-switching is a pervasive phenomenon in multilingual communication, yet the robustness of large language models (LLMs) in mixed-language settings remains insufficiently understood. In this work, we present a comprehe…