Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Observing that in these domains language changes are mostly associated with speaker changes, we propose a cascaded system consisting of speaker diarization and language identification and compare it with more traditional language identification and language diarization systems. Results show that the proposed system often achieves lower language classification and language diarization error rates (up to 10% relative language diarization error reduction and 60% relative language confusion reduction) and leads to lower WERs on multilingual test sets (more than 8% relative WER reduction), while at the same time does not negatively affect speech recognition on monolingual audio (with an absolute WER increase between 0.1% and 0.7% w.r.t. monolingual ASR).
Code (0)
등록된 구현이 없습니다.
Tasks
Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognitionSpeech RecognitionSpoken language identificationSimilar Papers 제목 키워드 기반
Multimodal Modeling For Spoken Language Identification
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have …
Language IdentificationSpoken language identificationDeep learning-based end-to-end spoken language identification system for domain-mismatched scenario
Domain mismatch is a critical issue when it comes to spoken language identification. To overcome the domain mismatch problem, we have applied several architectures and deep learning strategies which have shown good resul…
Language IdentificationSpeaker VerificationSpoken language identificationAutomatic Spoken Language Identification using a Time-Delay Neural Network
Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage Modeling+4Exploiting Spectral Augmentation for Code-Switched Spoken Language Identification
Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2Automatic Spoken Language Identification Utilizing Acoustic and Phonetic Speech Information
Automatic spoken Language Identification (LID) is the process of identifying the language spoken within an utterance. The challenge that this task presents is that no prior information is available indicating the content…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identification