Multi-resolution location-based training for multi-channel continuous speech separation
The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently, location-based training (LBT) was proposed as a new training criterion for multi-channel talker-independent speaker separation. Assuming fixed array geometry, LBT outperforms widely-used permutation-invariant training in fully overlapped utterances and matched reverberant conditions. This paper extends LBT to conversational multi-channel speaker separation. We introduce multi-resolution LBT to estimate the complex spectrograms from low to high time and frequency resolutions. With multi-resolution LBT, convolutional kernels are assigned consistently based on speaker locations in physical space. Evaluation results show that multi-resolution LBT consistently outperforms other competitive methods on the recorded LibriCSS corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Separationspeech-recognitionSpeech RecognitionSpeech SeparationSimilar Papers 제목 키워드 기반
Joint Channel Estimation and Mixed-ADCs Allocation for Massive MIMO via Deep Learning
Millimeter wave (mmWave) multi-user massive multi-input multi-output (MIMO) is a promising technique for the next generation communication systems. However, the hardware cost and power consumption grow significantly as t…
Location-aware Channel Estimation for RIS-aided mmWave MIMO Systems via Atomic Norm Minimization
In this paper, we propose a location-aware channel estimation based on the atomic norm minimization (ANM) for the reconfigurable intelligent surface (RIS)-aided millimeter-wave multiple-input-multiple-output (MIMO) syste…
Key Issues in Wireless Transmission for NTN-Assisted Internet of Things
Non-terrestrial networks (NTNs) have become appealing resolutions for seamless coverage in the next-generation wireless transmission, where a large number of Internet of Things (IoT) devices diversely distributed can be …
A Two-Stage Radar Sensing Approach based on MIMO-OFDM Technology
Recently, integrating the communication and sensing functions into a common network has attracted a great amount of attention. This paper considers the advanced signal processing techniques for enabling the radar to sens…
compressed sensingVocal Bursts Valence PredictionCoordinate-Queryable Neural Field Reconstruction for EEG Spatial Super-Resolution with Unseen-Electrode Generation
EEG spatial super-resolution (EEGSR) in real deployments is challenged by random channel missingness, unstable electrode quality, and changing visible-channel patterns caused by bad contacts or device variability. Most e…