paper-with-me

홈 › Papers

Online Binaural Speech Separation of Moving Speakers With a Wavesplit Network

2023-03-13 · Cong Han, Nima Mesgarani

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, however, the order of outputs can be inconsistent over time particularly in long-form speech separation. This situation which is referred to as the speaker swap problem is even more problematic when speakers constantly move in space and therefore poses a challenge for consistent placement of speakers in output channels. Here, we describe a real-time binaural speech separation model based on a Wavesplit network to mitigate the speaker swap problem for moving speaker separation. Our model computes a speaker embedding for each speaker at each time frame from the mixed audio, aggregates embeddings using online clustering, and uses cluster centroids as speaker profiles to track each speaker throughout the long duration. Experimental results on reverberant, long-form moving multitalker speech separation show that the proposed method is less prone to speaker swap and achieves comparable performance with u-PIT based models with ground truth tracking in both separation accuracy and preserving the interaural cues.

📄 PDF Abstract BibTeX arXiv:2303.07458

Code (0)

등록된 구현이 없습니다.

Tasks

Online ClusteringSpeaker SeparationSpeech Separation

Similar Papers 제목 키워드 기반

Spatial Speech Translation: Translating Across Space With Binaural Hearables

2025-04-25 · Tuochao Chen, Qirui Wang, Runlin He, Shyam Gollakota

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce …

blind source separationTranslation

Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

2025-09-16 · Manan Mittal, Thomas Deppisch, Joseph Forrer, Chris Le Sueur 외 arxiv

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to e…

Online Self-Attentive Gated RNNs for Real-Time Speaker Separation

2021-06-25 · Ori Kabeli, Yossi Adi, Zhenyu Tang, Buye Xu 외

Deep neural networks have recently shown great success in the task of blind source separation, both under monaural and binaural settings. Although these methods were shown to produce high-quality separations, they were m…

blind source separationSpeaker Separation

Online speaker diarization of meetings guided by speech separation

2024-01-30 · Elio Gruttadauria, Mathieu Fontaine, Slim Essid

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation mode…

Action DetectionActivity Detectionspeaker-diarizationSpeaker Diarization+1

The Cone of Silence: Speech Separation by Localization

2020-10-12 · NeurIPS 2020 12 · Teerapat Jenrungrot, Vivek Jayaram, Steve Seitz, Ira Kemelmacher-Shlizerman

Given a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the w…

Audio Source SeparationSpeech Separation