paper-with-me

Papers

SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation

2025-01-20 · Ziling Huang, Haixin Guan, Haoran Wei, Yanhua Long

Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired speech. However, these approaches suffer from significant model complexity and often underutilize enrollment speaker information, limiting the potential performance of the PSE model. To address these limitations, we propose a novel Speaker Encoder-Free PSE network, termed SEF-PNet, which fully exploits the information present in both the enrollment speech and noisy mixtures. SEF-PNet incorporates two key innovations: Interactive Speaker Adaptation (ISA) and Local-Global Context Aggregation (LCA). ISA dynamically modulates the interactions between enrollment and noisy signals to enhance the speaker adaptation, while LCA employs advanced channel attention within the PSE encoder to effectively integrate local and global contextual information, thus improving feature learning. Experiments on the Libri2Mix dataset demonstrate that SEF-PNet significantly outperforms baseline models, achieving state-of-the-art PSE performance.

📄 PDF Abstract BibTeX arXiv:2501.11274

Code (1)

ishuangziling/sef-pnet 공식 구현 pytorch

Tasks

Speaker VerificationSpeech Enhancement

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Personalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancement

2021-06-08 · Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik 외

The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present…

Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement

2025-05-18 · Ziling Huang, Haixin Guan, Yanhua Long

Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party…

DisentanglementSpeech Enhancement

Variational Autoencoder for Personalized Pathological Speech Enhancement

2025-03-18 · Mingchi Hou, Ina Kodrasi

The generalizability of speech enhancement (SE) models across speaker conditions remains largely unexplored, despite its critical importance for broader applicability. This paper investigates the performance of the hybri…

Speech Enhancement

EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers

2022-03-31 · Soumi Maiti, Yushi Ueda, Shinji Watanabe, Chunlei Zhang 외

In this paper, we present a novel framework that jointly performs three tasks: speaker diarization, speech separation, and speaker counting. Our proposed framework integrates speaker diarization based on end-to-end neura…

Decoderspeaker-diarizationSpeaker DiarizationSpeech Separation

Adversarial speech for voice privacy protection from Personalized Speech generation

2024-01-22 · Shihao Chen, Liping Chen, Jie Zhang, KongAik Lee 외

The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human list…

Speaker Verificationtext-to-speechText to SpeechVoice Conversion