Implicit spoken language diarization
Spoken language diarization (LD) and related tasks are mostly explored using the phonotactic approach. Phonotactic approaches mostly use explicit way of language modeling, hence requiring intermediate phoneme modeling and transcribed data. Alternatively, the ability of deep learning approaches to model temporal dynamics may help for the implicit modeling of language information through deep embedding vectors. Hence this work initially explores the available speaker diarization frameworks that capture speaker information implicitly to perform LD tasks. The performance of the LD system on synthetic code-switch data using the end-to-end x-vector approach is 6.78% and 7.06%, and for practical data is 22.50% and 60.38%, in terms of diarization error rate and Jaccard error rate (JER), respectively. The performance degradation is due to the data imbalance and resolved to some extent by using pre-trained wave2vec embeddings that provide a relative improvement of 30.74% in terms of JER.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModelingLanguage Modellingspeaker-diarizationSpeaker DiarizationSimilar Papers 제목 키워드 기반
Implicit Self-supervised Language Representation for Spoken Language Diarization
In a code-switched (CS) scenario, the use of spoken language diarization (LD) as a pre-possessing system is essential. Further, the use of implicit frameworks is preferable over the explicit framework, as it can be easil…
speaker-diarizationSpeaker DiarizationExploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech
This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Obser…
Language Identificationspeaker-diarizationSpeaker Diarizationspeech-recognition+2Towards Measuring and Scoring Speaker Diarization Fairness
Speaker diarization, or the task of finding "who spoke and when", is now used in almost every speech processing application. Nevertheless, its fairness has not yet been evaluated because there was no protocol to study it…
FairnessSentencespeaker-diarizationSpeaker DiarizationSAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
In this paper, we present a neural spoken language diarization model that supports an unconstrained span of languages within a single framework. Our approach integrates a learnable query-based architecture grounded in mu…
Data AugmentationImproving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and ofte…
speaker-diarizationSpeaker DiarizationSpoken Language Understanding