paper-with-me

홈 › Papers

A Survey of Deep Learning for Complex Speech Spectrograms

2025-05-13 · Yuying Xie, Zheng-Hua Tan

Recent advancements in deep learning have significantly impacted the field of speech signal processing, particularly in the analysis and manipulation of complex spectrograms. This survey provides a comprehensive overview of the state-of-the-art techniques leveraging deep neural networks for processing complex spectrograms, which encapsulate both magnitude and phase information. We begin by introducing complex spectrograms and their associated features for various speech processing tasks. Next, we explore the key components and architectures of complex-valued neural networks, which are specifically designed to handle complex-valued data and have been applied for complex spectrogram processing. We then discuss various training strategies and loss functions tailored for training neural networks to process and model complex spectrograms. The survey further examines key applications, including phase retrieval, speech enhancement, and speech separation, where deep learning has achieved significant progress by leveraging complex spectrograms or their derived feature representations. Additionally, we examine the intersection of complex spectrograms with generative models. This survey aims to serve as a valuable resource for researchers and practitioners in the field of speech signal processing and complex-valued neural networks.

📄 PDF Abstract BibTeX arXiv:2505.08694

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSpeech EnhancementSpeech SeparationSurvey

Similar Papers 제목 키워드 기반

Complex spectrogram enhancement by convolutional neural network with multi-metrics learning

2017-04-27 · Szu-Wei Fu, Ting-yao Hu, Yu Tsao, Xugang Lu

This paper aims to address two issues existing in the current speech enhancement methods: 1) the difficulty of phase estimations; 2) a single objective function cannot consider multiple metrics simultaneously. To solve t…

Speech Enhancement

High-quality Speech Synthesis Using Super-resolution Mel-Spectrogram

2019-12-03 · Leyuan Sheng, Dong-Yan Huang, Evgeniy N. Pavlovskiy

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized spe…

Image-to-Image TranslationSpeech EnhancementSpeech SynthesisSuper-Resolution+2

Deep Transform: Cocktail Party Source Separation via Complex Convolution in a Deep Neural Network

2015-04-12 · Andrew J. R. Simpson

Convolutional deep neural networks (DNN) are state of the art in many engineering problems but have not yet addressed the issue of how to deal with complex spectrograms. Here, we use circular statistics to provide a conv…

Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition

2024-10-03 · Olga Iakovenko, Ivan Bondarenko

For many Automatic Speech Recognition (ASR) tasks audio features as spectrograms show better results than Mel-frequency Cepstral Coefficients (MFCC), but in practice they are hard to use due to a complex dimensionality o…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Single channel speech enhancement by colored spectrograms

2023-10-26 · Sania Gul, Muhammad Salman Khan, Muhammad Fazeel

Speech enhancement concerns the processes required to remove unwanted background sounds from the target speech to improve its quality and intelligibility. In this paper, a novel approach for single-channel speech enhance…

DenoisingGenerative Adversarial NetworkSpeech Enhancement