paper-with-me

홈 › Papers

Improving Channel Decorrelation for Multi-Channel Target Speech Extraction

2021-06-06 · Jiangyu Han, Wei Rao, Yannan Wang, Yanhua Long

Target speech extraction has attracted widespread attention. When microphone arrays are available, the additional spatial information can be helpful in extracting the target speech. We have recently proposed a channel decorrelation (CD) mechanism to extract the inter-channel differential information to enhance the reference channel encoder representation. Although the proposed mechanism has shown promising results for extracting the target speech from mixtures, the extraction performance is still limited by the nature of the original decorrelation theory. In this paper, we propose two methods to broaden the horizon of the original channel decorrelation, by replacing the original softmax-based inter-channel similarity between encoder representations, using an unrolled probability and a normalized cosine-based similarity at the dimensional-level. Moreover, new combination strategies of the CD-based spatial information and target speaker adaptation of parallel encoder outputs are also investigated. Experiments on the reverberant WSJ0 2-mix show that the improved CD can result in more discriminative differential information and the new adaptation strategy is also very effective to improve the target speech extraction.

📄 PDF Abstract BibTeX arXiv:2106.03113

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Extraction

Similar Papers 제목 키워드 기반

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

2024-01-07 · Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 외

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is su…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

Enhancing Multivariate Time Series Forecasting with Mutual Information-driven Cross-Variable and Temporal Modeling

2024-03-01 · shiyi qi, Liangjian Wen, Yiduo Li, Yuanhang Yang 외

Recent advancements have underscored the impact of deep learning techniques on multivariate time series forecasting (MTSF). Generally, these techniques are bifurcated into two categories: Channel-independence and Channel…

Multivariate Time Series ForecastingTime SeriesTime Series Forecasting

Stochastic Channel Decorrelation Network and Its Application to Visual Tracking

2018-07-03 · Jie Guo, Tingfa Xu, Shenwang Jiang, Ziyi Shen

Deep convolutional neural networks (CNNs) have dominated many computer vision domains because of their great power to extract good features automatically. However, many deep CNNs-based computer vison tasks suffer from la…

DiversityVisual Tracking

ARiSE: Auto-Regressive Multi-Channel Speech Enhancement

2025-05-28 · Pengjie Shen, Xueliang Zhang, Zhong-Qiu Wang

We propose ARiSE, an auto-regressive algorithm for multi-channel speech enhancement. ARiSE improves existing deep neural network (DNN) based frame-online multi-channel speech enhancement models by introducing auto-regres…

Speech Enhancement

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition