paper-with-me

Papers

Investigating self-supervised features for expressive, multilingual voice conversion

2025-05-13 · Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer, Iván Vallés-Pérez, Roberto Barra-Chicote, Biel Tura-Vecino, Adam Gabryś, Piotr Bilinski, Thomas Merritt, Jaime Lorenzo-Trueba

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to produce. Unsupervised approaches are typically trained to reconstruct the input signal, which is composed of the content and the speaker information. Disentangling these components is a challenge and often leads to speaker leakage or prosodic information removal. In this paper, we explore voice conversion by leveraging the potential of self-supervised learning (SSL). A combination of the latent representations of SSL models, concatenated with speaker embeddings, is fed to a vocoder which is trained to reconstruct the input. Zero-shot voice conversion results show that this approach allows to keep the prosody and content of the source speaker while matching the speaker similarity of a VC system based on phonetic posteriorgrams (PPGs).

📄 PDF Abstract BibTeX arXiv:2505.08278

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeech SynthesisVoice Conversion

Similar Papers 제목 키워드 기반

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

2025-06-04 · Seymanur Aktı, Tuan Nam Nguyen, Alexander Waibel

Expressive voice conversion aims to transfer both speaker identity and expressive attributes from a target speech to a given source speech. In this work, we improve over a self-supervised, non-autoregressive framework wi…

DisentanglementStyle TransferVoice Conversion

Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision

2023-12-30 · Chih-Kai Yang, Kuan-Po Huang, Ke-Han Lu, Chun-Yi Kuan 외

This work evaluated several cutting-edge large-scale foundation models based on self-supervision or weak supervision, including SeamlessM4T, SeamlessM4T v2, and Whisper-large-v3, on three code-switched corpora. We found …

Speech-to-TextSpeech-to-Text Translation

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

2025-12-22 · Angelo Ortiz Tandazo, Manel Khentout, Youssef Benchekroun, Thomas Hueber 외 arxiv

This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representation learning. We continue HuBERT pre-training with supervision based on …

Self-Supervised LearningRepresentation Learning

Two Front-Ends, One Model : Fusing Heterogeneous Speech Features for Low Resource ASR with Multilingual Pre-Training

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Transfer learning is widely applied in various deep learning-based speech tasks, especially for tasks with a limited amount of data. Recent studies in transfer learning mainly focused on either supervised or self-supervi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

2022-03-28 · ACL 2022 5 · Zijie Huang, Zheng Li, Haoming Jiang, Tianyu Cao 외

Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various langua…

Knowledge Graph Completion