paper-with-me

Papers

Dynamically locating multiple speakers based on the time-frequency domain

2021-01-01 · Hodaya Hammer, Shlomo Chazan, Jacob Goldberger, Sharon Gannot

In this study we present a deep neural network-based online multi-speaker localisation algorithm based on a multi-microphone array. A fully convolutional network is trained with instantaneous spatial features to estimate the direction of arrival for each time-frequency bin. The high resolution classification enables the network to accurately and simultaneously localize and track multiple speakers, both static and dynamic. Elaborated experimental study using simulated and real-life recordings in static and dynamic scenarios, demonstrates that the proposed algorithm significantly outperforms both classic and recent deep-learning-based algorithms.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FCN Approach for Dynamically Locating Multiple Speakers

2020-08-26 · Hodaya Hammer, Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot

In this paper, we present a deep neural network-based online multi-speaker localisation algorithm. Following the W-disjoint orthogonality principle in the spectral domain, each time-frequency (TF) bin is dominated by a s…

Comparison of Frequency-Fusion Mechanisms for Binaural Direction-of-Arrival Estimation for Multiple Speakers

2024-01-15 · Daniel Fejgin, Elior Hadad, Sharon Gannot, Zbyněk Koldovský 외

To estimate the direction of arrival (DOA) of multiple speakers with methods that use prototype transfer functions, frequency-dependent spatial spectra (SPS) are usually constructed. To make the DOA estimation robust, SP…

Direction of Arrival Estimation

3D-Speaker: A Large-Scale Multi-Device, Multi-Distance, and Multi-Dialect Corpus for Speech Representation Disentanglement

2023-06-27 · Siqi Zheng, Luyao Cheng, Yafeng Chen, Hui Wang 외

Disentangling uncorrelated information in speech utterances is a crucial research topic within speech community. Different speech-related tasks focus on extracting distinct speech representations while minimizing the aff…

DisentanglementSelf-Supervised Learning

Coherence-Based Frequency Subset Selection For Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers

2022-05-18 · Daniel Fejgin, Simon Doclo

Recently, a method has been proposed to estimate the direction of arrival (DOA) of a single speaker by minimizing the frequency-averaged Hermitian angle between an estimated relative transfer function (RTF) vector and a …

Direction of Arrival Estimation

Non-native Accent Partitioning for Speakers of Indian Regional Languages

2019-12-01 · ICON 2019 12 · Radha Krishna Guntur, Krishnan Ramakrishnan, Vinay Kumar Mittal

Acoustic features extracted from the speech signal can help in identifying speaker related multiple information such as geographical origin, regional accent and nativity. In this paper, classification of native speakers …