paper-with-me

Papers

Speech collage: code-switched audio generation by collaging monolingual corpora

2023-09-27 · Amir Hussein, Dorsa Zeinali, Ondřej Klejch, Matthew Wiesner, Brian Yan, Shammur Chowdhury, Ahmed Ali, Shinji Watanabe, Sanjeev Khudanpur

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segments. We further improve the smoothness quality of audio generation using an overlap-add approach. We investigate the impact of generated data on speech recognition in two scenarios: using in-domain CS text and a zero-shot approach with synthesized CS text. Empirical results highlight up to 34.4% and 16.2% relative reductions in Mixed-Error Rate and Word-Error Rate for in-domain and zero-shot scenarios, respectively. Lastly, we demonstrate that CS augmentation bolsters the model's code-switching inclination and reduces its monolingual bias.

📄 PDF Abstract BibTeX arXiv:2309.15674

Code (1)

jsalt2022codeswitchingasr/generating-code-switched-audio 공식 구현 pytorch

Tasks

Audio GenerationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

2026-06-04 · Virginia Ceccatelli, Yejin Jeon, David Ifeoluwa Adelani arxiv

Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on monolingual, text-based harmful prompts. This leaves their generalizabi…

Reinforcement Learning for Data-Efficient Code-Switched ASR

2026-07-02 · Ziwei Ye, Peter Vickers arxiv

Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable…

Reinforcement Learning

AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR

2025-06-17 · Tuan Nguyen, Huy-Dat Tran

Developing code-switched ASR systems is challenging due to language ambiguity and limited exposure to multilingual, code-switched data, while collecting such speech is costly. Prior work generates synthetic audio from te…

Decoder

Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data

2024-09-17 · Jing Xu, Daxin Tan, Jiaqi Wang, Xiao Chen

While large language models (LLMs) have been explored in the speech domain for both generation and recognition tasks, their applications are predominantly confined to the monolingual scenario, with limited exploration in…

Speech Synthesis

Multi-Modal Transformers Utterance-Level Code-Switching Detection

2020-11-04 · Krishna D N

An utterance that contains speech from multiple languages is known as a code-switched sentence. In this work, we propose a novel technique to predict whether given audio is mono-lingual or code-switched. We propose a mul…

Sentence