paper-with-me

홈 › Papers

MBTFNet: Multi-Band Temporal-Frequency Neural Network For Singing Voice Enhancement

2023-10-06 · Weiming Xu, Zhouxuan Chen, Zhili Tan, Shubo Lv, Runduo Han, Wenjiang Zhou, Weifeng Zhao, Lei Xie

A typical neural speech enhancement (SE) approach mainly handles speech and noise mixtures, which is not optimal for singing voice enhancement scenarios. Music source separation (MSS) models treat vocals and various accompaniment components equally, which may reduce performance compared to the model that only considers vocal enhancement. In this paper, we propose a novel multi-band temporal-frequency neural network (MBTFNet) for singing voice enhancement, which particularly removes background music, noise and even backing vocals from singing recordings. MBTFNet combines inter and intra-band modeling for better processing of full-band signals. Dual-path modeling are introduced to expand the receptive field of the model. We propose an implicit personalized enhancement (IPE) stage based on signal-to-noise ratio (SNR) estimation, which further improves the performance of MBTFNet. Experiments show that our proposed model significantly outperforms several state-of-the-art SE and MSS models.

📄 PDF Abstract BibTeX arXiv:2310.04369

Code (0)

등록된 구현이 없습니다.

Tasks

Music Source SeparationSpeech Enhancement

Similar Papers 제목 키워드 기반

Multi-Band Multi-Resolution Fully Convolutional Neural Networks for Singing Voice Separation

2019-10-21 · Emad M. Grais, Fei Zhao, Mark D. Plumbley

Deep neural networks with convolutional layers usually process the entire spectrogram of an audio signal with the same time-frequency resolutions, number of filters, and dimensionality reduction scale. According to the c…

Dimensionality Reduction

HiFiSinger: Towards High-Fidelity Neural Singing Voice Synthesis

2020-09-03 · Jiawei Chen, Xu Tan, Jian Luan, Tao Qin 외

High-fidelity singing voices usually require higher sampling rate (e.g., 48kHz) to convey expression and emotion. However, higher sampling rate causes the wider frequency band and longer waveform sequences and throws cha…

Singing Voice SynthesisVocal Bursts Intensity Prediction

Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network

2022-10-26 · Interspeech 2023 8 · Chunhui Wang, Chang Zeng, Xing He

XiaoiceSing is a singing voice synthesis (SVS) system that aims at generating 48kHz singing voices. However, the mel-spectrogram generated by it is over-smoothing in middle- and high-frequency areas due to no special des…

Generative Adversarial NetworkSinging Voice Synthesis

SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice Generation

2021-10-14 · Rongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren 외

Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts…

Generative Adversarial NetworkGPUSinging Voice SynthesisSpeech Synthesis+3

Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus

2021-12-20 · MM '21: Proceedings of the 29th ACM International Conference on Multimedia 2021 10 · Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu 외

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not me…

Audio GenerationSinging Voice SynthesisText-To-Speech Synthesis