Automatic Spoken Language Identification using a Time-Delay Neural Network
Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained to distinguish between Arabic, Spanish, French, and Turkish based on nothing more than recorded speech. A pre-existing multilingual dataset was used to train a series of acoustic models based on the Tedlium TDNN model to perform automatic speech recognition. The system was provided with a custom multilingual language model and a specialized pronunciation lexicon with language names prepended to phones. The trained model was used to generate phone alignments to test data from all four languages, and languages were predicted based on a voting scheme choosing the most common language prepend in an utterance. Accuracy was measured by comparing predicted languages to known languages, and was determined to be very high in identifying Spanish and Arabic, and somewhat lower in identifying Turkish and French.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpoken language identificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Multimodal Modeling For Spoken Language Identification
Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have …
Language IdentificationSpoken language identificationExploiting Spectral Augmentation for Code-Switched Spoken Language Identification
Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2Automatic Spoken Language Identification Utilizing Acoustic and Phonetic Speech Information
Automatic spoken Language Identification (LID) is the process of identifying the language spoken within an utterance. The challenge that this task presents is that no prior information is available indicating the content…
Language Identificationspeech-recognitionSpeech RecognitionSpoken language identificationImproving Multilingual ASR in the Wild Using Simple N-best Re-ranking
Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical setting…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language IdentificationRe-Ranking+3Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…
Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2