paper-with-me

홈 › Papers

DPATD: Dual-Phase Audio Transformer for Denoising

2023-10-30 · Junhui Li, Pu Wang, Jialu Li, Xinzhe Wang, Youshan Zhang

Recent high-performance transformer-based speech enhancement models demonstrate that time domain methods could achieve similar performance as time-frequency domain methods. However, time-domain speech enhancement systems typically receive input audio sequences consisting of a large number of time steps, making it challenging to model extremely long sequences and train models to perform adequately. In this paper, we utilize smaller audio chunks as input to achieve efficient utilization of audio information to address the above challenges. We propose a dual-phase audio transformer for denoising (DPATD), a novel model to organize transformer layers in a deep structure to learn clean audio sequences for denoising. DPATD splits the audio input into smaller chunks, where the input length can be proportional to the square root of the original sequence length. Our memory-compressed explainable attention is efficient and converges faster compared to the frequently used self-attention module. Extensive experiments demonstrate that our model outperforms state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2310.19588

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

Complex Image Generation SwinTransformer Network for Audio Denoising

2023-10-24 · Youshan Zhang, Jialu Li

Achieving high-performance audio denoising is still a challenging task in real-world applications. Existing time-frequency methods often ignore the quality of generated frequency domain images. This paper converts the au…

Audio DenoisingDenoisingImage Generation

Self-Supervised Audio-and-Text Pre-training with Extremely Low-Resource Parallel Data

2022-04-10 · Yu Kang, Tianqiao Liu, Hang Li, Yang Hao 외

Multimodal pre-training for audio-and-text has recently been proved to be effective and has significantly improved the performance of many downstream speech understanding tasks. However, these state-of-the-art pre-traini…

Denoising

Low-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing

2019-03-13

Low-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their ampl…

Audio DenoisingAudio Signal ProcessingDenoising

Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators

2024-08-11 · Yifan Pu, Zhuofan Xia, Jiayi Guo, Dongchen Han 외

This paper identifies significant redundancy in the query-key interactions within self-attention mechanisms of diffusion transformer models, particularly during the early stages of denoising diffusion steps. In response …

Denoising

Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition

2026-03-22 · Ying Liu, Yuntao Shou, Wei Ai, Tao Meng 외 arxiv

In real-world scenarios, audio and video signals are often subject to environmental noise and limited acquisition conditions, resulting in extracted features containing excessive noise. Furthermore, there is an imbalance…

Emotion Recognition