paper-with-me

Papers

Progressive Residual Extraction based Pre-training for Speech Representation Learning

2024-08-31 · Tianrui Wang, Jin Li, Ziyang Ma, Rui Cao, Xie Chen, Longbiao Wang, Meng Ge, Xiaobao Wang, Yuguang Wang, Jianwu Dang, Nyima Tashi

Self-supervised learning (SSL) has garnered significant attention in speech processing, excelling in linguistic tasks such as speech recognition. However, jointly improving the performance of pre-trained models on various downstream tasks, each requiring different speech information, poses significant challenges. To this purpose, we propose a progressive residual extraction based self-supervised learning method, named ProgRE. Specifically, we introduce two lightweight and specialized task modules into an encoder-style SSL backbone to enhance its ability to extract pitch variation and speaker information from speech. Furthermore, to prevent the interference of reinforced pitch variation and speaker information with irrelevant content information learning, we residually remove the information extracted by these two modules from the main branch. The main branch is then trained using HuBERT's speech masking prediction to ensure the performance of the Transformer's deep-layer features on content tasks. In this way, we can progressively extract pitch variation, speaker, and content representations from the input speech. Finally, we can combine multiple representations with diverse speech information using different layer weights to obtain task-specific representations for various downstream tasks. Experimental results indicate that our proposed method achieves joint performance improvements on various tasks, such as speaker identification, speech recognition, emotion recognition, speech enhancement, and voice conversion, compared to excellent SSL methods such as wav2vec2.0, HuBERT, and WavLM.

📄 PDF Abstract BibTeX arXiv:2409.00387

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionRepresentation LearningSelf-Supervised LearningSpeaker IdentificationSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech Representation LearningVoice Conversion

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

StepAudio 3 Gen Technical Report

2026-09-11 · Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang 외 hf

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types…

Audio Generation

RESAR-BEV: An Explainable Progressive Residual Autoregressive Approach for Camera-Radar Fusion in BEV Segmentation

2025-05-10 · Zhiwen Zeng, Yunfei Yin, Zheng Yuan, Argho Dey 외

Bird's-Eye-View (BEV) semantic segmentation provides comprehensive environmental perception for autonomous driving but suffers multi-modal misalignment and sensor noise. We propose RESAR-BEV, a progressive refinement fra…

Autonomous DrivingBEV SegmentationSemantic Segmentation

APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

2025-04-29 · Zhicheng Lian, Lizhi Wang, Hua Huang

Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have …

Quantization

Progressive Approximation in Deep Residual Networks: Theory and Validation

2026-04-27 · Wei Wang, Xiao-Yong Wei, Qing Li arxiv

The Universal Approximation Theorem (UAT) guarantees universal function approximation but does not explain how residual models distribute approximation across layers. We reframe residual networks as a layer-wise approxim…

Representation LearningImage Classification

TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 DNS Challenge

2023-03-14 · Yukai Ju, Jun Chen, Shimin Zhang, Shulin He 외

This paper introduces the Unbeatable Team's submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version -- TEA-PSE 3.0. Specifically, TEA-PSE 3.0 in…

Speech Enhancement