paper-with-me

Papers

FCN Approach for Dynamically Locating Multiple Speakers

2020-08-26 · Hodaya Hammer, Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot

In this paper, we present a deep neural network-based online multi-speaker localisation algorithm. Following the W-disjoint orthogonality principle in the spectral domain, each time-frequency (TF) bin is dominated by a single speaker, and hence by a single direction of arrival (DOA). A fully convolutional network is trained with instantaneous spatial features to estimate the DOA for each TF bin. The high resolution classification enables the network to accurately and simultaneously localize and track multiple speakers, both static and dynamic. Elaborated experimental study using both simulated and real-life recordings in static and dynamic scenarios, confirms that the proposed algorithm outperforms both classic and recent deep-learning-based algorithms.

📄 PDF Abstract BibTeX arXiv:2008.11845

Code (1)

MindSpore-paper-code-2/code399/tree/main/fcn-4 mindspore

Similar Papers 제목 키워드 기반

Dynamically locating multiple speakers based on the time-frequency domain

2021-01-01 · Hodaya Hammer, Shlomo Chazan, Jacob Goldberger, Sharon Gannot

In this study we present a deep neural network-based online multi-speaker localisation algorithm based on a multi-microphone array. A fully convolutional network is trained with instantaneous spatial features to estimate…

3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

2023-06-27 · Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang 외

Disentangling uncorrelated information in speech utterances is a crucial research topic within speech community. Different speech-related tasks focus on extracting distinct speech representations while minimizing the aff…

DisentanglementSelf-Supervised Learning

Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers

2025-05-22 · Yuzhu Wang, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously perf…

Speech Separation

The Polish Vocabulary Size Test: A Novel Adaptive Test for Receptive Vocabulary Assessment

2025-07-26 · Danil Fokin, Monika Płużyczka, Grigory Golovin arxiv

We present the Polish Vocabulary Size Test (PVST), a novel tool for assessing the receptive vocabulary size of both native and non-native Polish speakers. Based on Item Response Theory and Computerized Adaptive Testing, …

Voice Separation with an Unknown Number of Multiple Speakers

2020-02-29 · ICML 2020 1 · Eliya Nachmani, Yossi Adi, Lior Wolf

We present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing st…

Speech Separation