paper-with-me

Papers

Reprogramming Self-supervised Learning-based Speech Representations for Speaker Anonymization

2023-11-17 · Xiaojiao Chen, Sheng Li, Jiyi Li, Hao Huang, Yang Cao, Liang He

Current speaker anonymization methods, especially with self-supervised learning (SSL) models, require massive computational resources when hiding speaker identity. This paper proposes an effective and parameter-efficient speaker anonymization method based on recent End-to-End model reprogramming technology. To improve the anonymization performance, we first extract speaker representation from large SSL models as the speaker identifies. To hide the speaker's identity, we reprogram the speaker representation by adapting the speaker to a pseudo domain. Extensive experiments are carried out on the VoicePrivacy Challenge (VPC) 2022 datasets to demonstrate the effectiveness of our proposed parameter-efficient learning anonymization methods. Additionally, while achieving comparable performance with the VPC 2022 strong baseline 1.b, our approach consumes less computational resources during anonymization.

📄 PDF Abstract BibTeX arXiv:2311.10664

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeaker anonymization

Similar Papers 제목 키워드 기반

Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering

2023-05-18 · Heng-Jui Chang, Alexander H. Liu, James Glass

Self-supervised speech representation models have succeeded in various tasks, but improving them for content-related problems using unlabeled data is challenging. We propose speaker-invariant clustering (Spin), a novel s…

Acoustic Unit DiscoveryClusteringGPUSelf-Supervised Learning+2

Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation

2025-05-25 · Giuseppe Ruggiero, Matteo Testa, Jurgen Van de Walle, Luigi di Caro

Self-supervised learning (SSL) has reduced the reliance on expensive labeling in speech technologies by learning meaningful representations from unannotated data. Since most SSL-based downstream tasks prioritize content …

DisentanglementSelf-Supervised LearningVoice Conversion

ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations

2023-03-01 · Neil Shah, Saiteja Kosgi, Vishal Tambrahalli, Neha Sahipjohn 외

We present ParrotTTS, a modularized text-to-speech synthesis model leveraging disentangled self-supervised speech representations. It can train a multi-speaker variant effectively using transcripts from a single speaker.…

Self-Supervised LearningSpeech Synthesistext-to-speechText to Speech+1

Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer

2023-09-14 · Yongqi Wang, Jionghao Bai, Rongjie Huang, RuiQi Li 외

Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech. Meanwhile, the scarcity of hig…

In-Context LearningLanguage ModelingLanguage ModellingSpeech-to-Speech Translation+2

ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers

2022-04-20 · Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni 외

Self-supervised learning in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of th…

DisentanglementSelf-Supervised Learning