paper-with-me

홈 › Papers

BASEN: Time-Domain Brain-Assisted Speech Enhancement Network with Convolutional Cross Attention in Multi-talker Conditions

2023-05-17 · Jie Zhang, Qing-Tian Xu, Qiu-Shi Zhu, Zhen-Hua Ling

Time-domain single-channel speech enhancement (SE) still remains challenging to extract the target speaker without any prior information on multi-talker conditions. It has been shown via auditory attention decoding that the brain activity of the listener contains the auditory information of the attended speaker. In this paper, we thus propose a novel time-domain brain-assisted SE network (BASEN) incorporating electroencephalography (EEG) signals recorded from the listener for extracting the target speaker from monaural speech mixtures. The proposed BASEN is based on the fully-convolutional time-domain audio separation network. In order to fully leverage the complementary information contained in the EEG signals, we further propose a convolutional multi-layer cross attention module to fuse the dual-branch features. Experimental results on a public dataset show that the proposed model outperforms the state-of-the-art method in several evaluation metrics. The reproducible code is available at https://github.com/jzhangU/Basen.git.

📄 PDF Abstract BibTeX arXiv:2305.09994

Code (1)

jzhangu/basen 공식 구현

Tasks

EEGSpeech Enhancement

Similar Papers 제목 키워드 기반

Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement

2024-09-19 · Keying Zuo, Qingtian Xu, Jie Zhang, ZhenHua Ling

Brain-assisted speech enhancement (BASE) aims to extract the target speaker in complex multi-talker scenarios using electroencephalogram (EEG) signals as an assistive modality, as the auditory attention of the listener c…

channel selectionEEGElectroencephalogram (EEG)Speech Enhancement

Sparsity-Driven EEG Channel Selection for Brain-Assisted Speech Enhancement

2023-11-22 · Jie Zhang, Qing-Tian Xu, Zhen-Hua Ling, Haizhou Li

Speech enhancement is widely used as a front-end to improve the speech quality in many audio systems, while it is hard to extract the target speech in multi-talker conditions without prior information on the speaker iden…

channel selectionEEGElectroencephalogram (EEG)Speech Enhancement

BASENet: Band-Adapted Speech Enhancement Network with Cross-Band Attention

2026-06-10 · Damien Martins Gomes, François Capman arxiv

Speech enhancement models typically apply uniform capacity across all frequencies, disregarding the non-uniform spectral resolution of human hearing. We propose BASENet, a frequency-adapted architecture that partitions t…

Speech Enhancement

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

2025-05-31 · Cunhang Fan, Ying Chen, Jian Zhou, Zexu Pan 외

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlo…

Contrastive LearningEEGTarget Speaker Extraction

Magnetoencephalography (MEG) Based Non-Invasive Chinese Speech Decoding

2025-06-15 · Zhihong Jia, Hongbin Wang, Yuanzhong Shen, Feng Hu 외

As an emerging paradigm of brain-computer interfaces (BCIs), speech BCI has the potential to directly reflect auditory perception and thoughts, offering a promising communication alternative for patients with aphasia. Ch…