paper-with-me

홈 › Papers

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model

2025-05-20 · Yang Xiang, Canan Huang, Desheng Hu, Jingguang Tian, Xinhui Hu, Chao Zhang

Most current speech enhancement (SE) methods recover clean speech from noisy inputs by directly estimating time-frequency masks or spectrums. However, these approaches often neglect the distinct attributes, such as semantic content and acoustic details, inherent in speech signals, which can hinder performance in downstream tasks. Moreover, their effectiveness tends to degrade in complex acoustic environments. To overcome these challenges, we propose a novel, semantic information-based, step-by-step factorized SE method using factorized codec and diffusion model. Unlike traditional SE methods, our hierarchical modeling of semantic and acoustic attributes enables more robust clean speech recovery, particularly in challenging acoustic scenarios. Moreover, this method offers further advantages for downstream TTS tasks. Experimental results demonstrate that our algorithm not only outperforms SOTA baselines in terms of speech quality but also enhances TTS performance in noisy environments.

📄 PDF Abstract BibTeX arXiv:2505.13843

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Factorized RVQ-GAN For Disentangled Speech Tokenization

2025-06-18 · Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos 외

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge dist…

DisentanglementKnowledge DistillationSpeech Tokenization

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

2025-02-05 · Jixun Yao, Hexin Liu, Chen Chen, Yuchen Hu 외

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns…

Language ModelingLanguage ModellingSpeech Enhancement

Speech Disorder Classification Using Extended Factorized Hierarchical Variational Auto-encoders

2021-06-14 · Jinzi Qi, Hugo Van hamme

Objective speech disorder classification for speakers with communication difficulty is desirable for diagnosis and administering therapy. With the current state of speech technology, it is evident to propose neural netwo…

ClassificationDisentanglementRepresentation LearningSentence

Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data

2017-09-22 · NeurIPS 2017 12 · Wei-Ning Hsu, Yu Zhang, James Glass

We present a factorized hierarchical variational autoencoder, which learns disentangled and interpretable representations from sequential data without supervision. Specifically, we exploit the multi-scale nature of infor…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speaker Verificationspeech-recognition+1

Disentangled Speech Representation Learning Based on Factorized Hierarchical Variational Autoencoder with Self-Supervised Objective

2022-04-05 · Yuying Xie, Thomas Arildsen, Zheng-Hua Tan

Disentangled representation learning aims to extract explanatory features or factors and retain salient information. Factorized hierarchical variational autoencoder (FHVAE) presents a way to disentangle a speech signal i…

DisentanglementRepresentation LearningSpeaker Recognitionspeech-recognition+3