paper-with-me

홈 › Papers

Stage-Wise and Prior-Aware Neural Speech Phase Prediction

2024-10-07 · Fei Liu, Yang Ai, Hui-Peng Du, Ye-Xin Lu, Rui-Chen Zheng, Zhen-Hua Ling

This paper proposes a novel Stage-wise and Prior-aware Neural Speech Phase Prediction (SP-NSPP) model, which predicts the phase spectrum from input amplitude spectrum by two-stage neural networks. In the initial prior-construction stage, we preliminarily predict a rough prior phase spectrum from the amplitude spectrum. The subsequent refinement stage transforms the amplitude spectrum into a refined high-quality phase spectrum conditioned on the prior phase. Networks in both stages use ConvNeXt v2 blocks as the backbone and adopt adversarial training by innovatively introducing a phase spectrum discriminator (PSD). To further improve the continuity of the refined phase, we also incorporate a time-frequency integrated difference (TFID) loss in the refinement stage. Experimental results confirm that, compared to neural network-based no-prior phase prediction methods, the proposed SP-NSPP achieves higher phase prediction accuracy, thanks to introducing the coarse phase priors and diverse training criteria. Compared to iterative phase estimation algorithms, our proposed SP-NSPP does not require multiple rounds of staged iterations, resulting in higher generation efficiency.

📄 PDF Abstract BibTeX arXiv:2410.04990

Code (0)

등록된 구현이 없습니다.

Tasks

Prediction

Methods 이 논문이 사용한 방법론

ConvNeXt 설명 없음

Similar Papers 제목 키워드 기반

Phase-aware Single-stage Speech Denoising and Dereverberation with U-Net

2020-06-01 · Interspeech 2020 6

In this work, we tackle a denoising and dereverberation problem with a single-stage framework. Although denoising and dereverberation may be considered two separate challenging tasks, and thus, two modules are typically …

DenoisingSpeech DenoisingSpeech Enhancement

Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

2024-06-17 · Kubilay Can Demir, Belen Lojo Rodriguez, Tobias Weise, Andreas Maier 외

To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a mu…

Surgical phase recognition

Phase-aware Speech Enhancement with Deep Complex U-Net

2019-03-07 · ICLR 2019 5 · Hyeong-Seok Choi, Jang-Hyun Kim, Jaesung Huh, Adrian Kim 외

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimat…

Speech Enhancementvalid

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

2026-06-26 · Stefano Giacomelli, Stefano Damiano, Claudia Rinaldi, Fabio Graziosi 외 arxiv

This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrain…

Sound Event Localization and DetectionNeural Architecture SearchAudio Tagging

SSR: Alignment-Aware Modality Connector for Speech Language Models

2024-09-30 · Weiting Tan, Hirofumi Inaguma, Ning Dong, Paden Tomasello 외

Fusing speech into pre-trained language model (SpeechLM) usually suffers from inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality. We propose SSR-Connector (Segmented Speech …

Language ModelingLanguage ModellingMMLU