paper-with-me

Papers

Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions

2024-06-12 · Anfeng Xu, Kevin Huang, Tiantian Feng, Lue Shen, Helen Tager-Flusberg, Shrikanth Narayanan

Speech foundation models, trained on vast datasets, have opened unique opportunities in addressing challenging low-resource speech understanding, such as child speech. In this work, we explore the capabilities of speech foundation models on child-adult speaker diarization. We show that exemplary foundation models can achieve 39.5% and 62.3% relative reductions in Diarization Error Rate and Speaker Confusion Rate, respectively, compared to previous speaker diarization methods. In addition, we benchmark and evaluate the speaker diarization results of the speech foundation models with varying the input audio window size, speaker demographics, and training data ratio. Our results highlight promising pathways for understanding and adopting speech foundation models to facilitate child speech understanding.

📄 PDF Abstract BibTeX arXiv:2406.07890

Code (1)

usc-sail/child-adult-diarization 공식 구현 pytorch

Tasks

speaker-diarizationSpeaker Diarization

Similar Papers 제목 키워드 기반

Data Efficient Child-Adult Speaker Diarization with Simulated Conversations

2024-09-13 · Anfeng Xu, Tiantian Feng, Helen Tager-Flusberg, Catherine Lord 외

Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, pub…

speaker-diarizationSpeaker DiarizationTransfer Learning

Ultrasound tongue imaging for diarization and alignment of child speech therapy sessions

2019-07-01 · Manuel Sam Ribeiro, Aciel Eshky, Korin Richmond, Steve Renals

We investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of…

speaker-diarizationSpeaker DiarizationWord Alignment

Multi-Stage Speaker Diarization for Noisy Classrooms

2025-05-16 · Ali Sartaz Khan, Tolulope Ogunremi, Ahmed Adel Attia, Dorottya Demszky

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording q…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

The Second DIHARD Diarization Challenge: Dataset, task, and baselines

2019-06-18 · Neville Ryant, Kenneth Church, Christopher Cieri, Alejandrina Cristia 외

This paper introduces the second DIHARD challenge, the second in a series of speaker diarization challenges intended to improve the robustness of diarization systems to variation in recording equipment, noise conditions,…

Action DetectionActivity DetectionLanguage AcquisitionSegmentation+3

Speaker diarization using latent space clustering in generative adversarial network

2019-10-24

In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) backprojection with the help of an encoder network. The proposed diarization system is trained join…

ClusteringDiagnosticGenerative Adversarial Networkspeaker-diarization+1