paper-with-me

Papers

Speaker Diarization and Identification from Single-Channel Classroom Audio Recording Using Virtual Microphones

2022-07-01 · Antonio Gomez

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking at the same time. To solve the problem, we assume the use of a single microphone per student group without any access to previous large datasets for training. This dissertation proposes a method of speaker identification using cross-correlation patterns associated to an array of virtual microphones, centered around the physical microphone. The virtual microphones are simulated by using approximate speaker geometry observed from a video recording. The patterns are constructed based on estimates of the room impulse responses for each virtual microphone. The correlation patterns are then used to identify the speakers. The proposed method is validated with classroom audios and shown to substantially outperform diarization services provided by Google Cloud and Amazon AWS.

📄 PDF Abstract BibTeX arXiv:2207.00660

Code (0)

등록된 구현이 없습니다.

Tasks

speaker-diarizationSpeaker DiarizationSpeaker Identification

Similar Papers 제목 키워드 기반

Multi-Stage Speaker Diarization for Noisy Classrooms

2025-05-16 · Ali Sartaz Khan, Tolulope Ogunremi, Ahmed Adel Attia, Dorottya Demszky

Speaker diarization, the process of identifying "who spoke when" in audio recordings, is essential for understanding classroom dynamics. However, classroom settings present distinct challenges, including poor recording q…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+5

Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization

2022-10-07 · Shota Horiguchi, Yuki Takashima, Shinji Watanabe, Paola Garcia

Due to the high performance of multi-channel speech processing, we can use the outputs from a multi-channel model as teacher labels when training a single-channel model with knowledge distillation. To the contrary, it is…

Knowledge Distillationspeaker-diarizationSpeaker DiarizationTransfer Learning

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

2021-05-05 · Soumi Maiti, Hakan Erdogan, Kevin Wilson, Scott Wisdom 외

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling speaker overlap and enabling straightforwar…

ClusteringSpeaker IdentificationTransfer Learning

Multi-Channel Sequence-to-Sequence Neural Diarization: Experimental Results for The MISP 2025 Challenge

2025-05-22 · Ming Cheng, Fei Su, Cancan Li, Juan Liu 외

This paper describes the speaker diarization system developed for the Multimodal Information-Based Speech Processing (MISP) 2025 Challenge. First, we utilize the Sequence-to-Sequence Neural Diarization (S2SND) framework …

speaker-diarizationSpeaker Diarization

Speaker Diarization: Using Recurrent Neural Networks

2020-06-10

Speaker Diarization is the problem of separating speakers in an audio. There could be any number of speakers and final result should state when speaker starts and ends. In this project, we analyze given audio file with 2…

speaker-diarizationSpeaker Diarization