paper-with-me

Papers

Speech Enhancement using Self-Adaptation and Multi-Head Self-Attention

2020-02-14 · Yuma Koizumi, Kohei Yatabe, Marc Delcroix, Yoshiki Masuyama, Daiki Takeuchi

This paper investigates a self-adaptation method for speech enhancement using auxiliary speaker-aware features; we extract a speaker representation used for adaptation directly from the test utterance. Conventional studies of deep neural network (DNN)--based speech enhancement mainly focus on building a speaker independent model. Meanwhile, in speech applications including speech recognition and synthesis, it is known that model adaptation to the target speaker improves the accuracy. Our research question is whether a DNN for speech enhancement can be adopted to unknown speakers without any auxiliary guidance signal in test-phase. To achieve this, we adopt multi-task learning of speech enhancement and speaker identification, and use the output of the final hidden layer of speaker identification branch as an auxiliary feature. In addition, we use multi-head self-attention for capturing long-term dependencies in the speech and noise. Experimental results on a public dataset show that our strategy achieves the state-of-the-art performance and also outperform conventional methods in terms of subjective quality.

📄 PDF Abstract BibTeX arXiv:2002.05873

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningSpeaker IdentificationSpeech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

U-Former: Improving Monaural Speech Enhancement with Multi-head Self and Cross Attention

2022-05-18 · Xinmeng Xu, Jianjun Hao

For supervised speech enhancement, contextual information is important for accurate spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term c…

DecoderSpeech Enhancement

Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments

2022-07-15 · Yicheng Du, Aditya Arie Nugraha, Kouhei Sekiguchi, Yoshiaki Bando 외

This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simula…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Distant Speech RecognitionNoisy Speech Recognition+3

MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement

2025-07-01 · Nikolai Lund Kühne, Jesper Jensen, Jan Østergaard, Zheng-Hua Tan

With the advent of new sequence models like Mamba and xLSTM, several studies have shown that these models match or outperform state-of-the-art models in single-channel speech enhancement, automatic speech recognition, an…

Automatic Speech RecognitionMambaSpeech Enhancementspeech-recognition+1

Continual self-training with bootstrapped remixing for speech enhancement

2021-10-19 · Efthymios Tzinis, Yossi Adi, Vamsi K. Ithapu, Buye Xu 외

We propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies includin…

Domain AdaptationSpeech EnhancementUnsupervised Domain Adaptation

Test-Time Training for Speech Enhancement

2025-08-03 · Avishkar Behera, Riya Ann Easow, Venkatesh Parvathala, K. Sri Rama Murty arxiv

This paper introduces a novel application of Test-Time Training (TTT) for Speech Enhancement, addressing the challenges posed by unpredictable noise conditions and domain shifts. This method combines a main speech enhanc…

Speech Enhancement