Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available.
Code (1)
Tasks
speaker-diarizationSpeaker DiarizationTransfer LearningSimilar Papers 제목 키워드 기반
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
Speech foundation models, trained on vast datasets, have opened unique opportunities in addressing challenging low-resource speech understanding, such as child speech. In this work, we explore the capabilities of speech …
speaker-diarizationSpeaker DiarizationMeta-learning for robust child-adult classification from speech
Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and interven…
ClassificationMeta-Learningspeaker-diarizationSpeaker Diarization+1Speaker Verification Experiments for Adults and Children Using Shared Embedding Spaces
For children, the system trained on a large corpus of adult speakers performed worse than a system trained on a much smaller corpus of children’s speech. This is due to the acoustic mismatch between training and testing …
Speaker VerificationEnhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
The assessment of children at risk of autism typically involves a clinician observing, taking notes, and rating children's behaviors. A machine learning model that can label adult and child audio may largely save labor i…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Phoneme RecognitionSelf-Supervised Learning+4Ultrasound tongue imaging for diarization and alignment of child speech therapy sessions
We investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of…
speaker-diarizationSpeaker DiarizationWord Alignment