paper-with-me

Papers

Towards Zero-Shot Code-Switched Speech Recognition

2022-11-02 · Brian Yan, Matthew Wiesner, Ondrej Klejch, Preethi Jyothi, Shinji Watanabe

In this work, we seek to build effective code-switched (CS) automatic speech recognition systems (ASR) under the zero-shot setting where no transcribed CS speech data is available for training. Previously proposed frameworks which conditionally factorize the bilingual task into its constituent monolingual parts are a promising starting point for leveraging monolingual data efficiently. However, these methods require the monolingual modules to perform language segmentation. That is, each monolingual module has to simultaneously detect CS points and transcribe speech segments of one language while ignoring those of other languages -- not a trivial task. We propose to simplify each monolingual module by allowing them to transcribe all speech segments indiscriminately with a monolingual script (i.e. transliteration). This simple modification passes the responsibility of CS point detection to subsequent bilingual modules which determine the final output by considering multiple monolingual transliterations along with external language model information. We apply this transliteration-based approach in an end-to-end differentiable neural network and demonstrate its efficacy for zero-shot CS ASR on Mandarin-English SEAME test sets.

📄 PDF Abstract BibTeX arXiv:2211.01458

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modellingspeech-recognitionSpeech RecognitionTransliteration

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Speech collage: code-switched audio generation by collaging monolingual corpora

2023-09-27 · Amir Hussein, Dorsa Zeinali, Ondřej Klejch, Matthew Wiesner 외

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a …

Audio GenerationAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1

Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization

2023-05-18 · Puyuan Peng, Brian Yan, Shinji Watanabe, David Harwath

We investigate the emergent abilities of the recently proposed web-scale speech model Whisper, by adapting it to unseen tasks with prompt engineering. We selected three tasks: audio-visual speech recognition (AVSR), code…

Audio-Visual Speech RecognitionPrompt Engineeringspeech-recognitionSpeech Recognition+1

Reinforcement Learning for Data-Efficient Code-Switched ASR

2026-07-02 · Ziwei Ye, Peter Vickers arxiv

Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable…

Reinforcement Learning

Learning not to Discriminate: Task Agnostic Learning for Improving Monolingual and Code-switched Speech Recognition

2020-06-09 · Gurunath Reddy Madhumani, Sanket Shah, Basil Abraham, Vikas Joshi 외

Recognizing code-switched speech is challenging for Automatic Speech Recognition (ASR) for a variety of reasons, including the lack of code-switched training data. Recently, we showed that monolingual ASR systems fine-tu…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Learning to Recognize Code-switched Speech Without Forgetting Monolingual Speech Recognition

2020-06-01 · Sanket Shah, Basil Abraham, Gurunath Reddy M, Sunayana Sitaram 외

Recently, there has been significant progress made in Automatic Speech Recognition (ASR) of code-switched speech, leading to gains in accuracy on code-switched datasets in many language pairs. Code-switched speech co-occ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition