paper-with-me

Papers

Disentangled Speech Embeddings using Cross-modal Self-supervision

2020-02-20 · Arsha Nagrani, Joon Son Chung, Samuel Albanie, Andrew Zisserman

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces and audio in video. The key idea behind our approach is to tease apart--without annotation--the representations of linguistic content and speaker identity. We construct a two-stream architecture which: (1) shares low-level features common to both representations; and (2) provides a natural mechanism for explicitly disentangling these factors, offering the potential for greater generalisation to novel combinations of content and identity and ultimately producing speaker identity representations that are more robust. We train our method on a large-scale audio-visual dataset of talking heads `in the wild', and demonstrate its efficacy by evaluating the learned speaker representations for standard speaker recognition performance.

📄 PDF Abstract BibTeX arXiv:2002.08742

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised LearningSpeaker Recognition

Similar Papers 제목 키워드 기반

Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model

2023-04-24 · Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya 외

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector…

RhythmSelf-Supervised LearningSpeech Synthesistext-to-speech+2

Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction

2023-12-16 · Zhaoxi Mu, Xinyu Yang, Sining Sun, Qing Yang

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic…

DisentanglementRepresentation LearningSpeech Extraction

ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech

2026-01-18 · Haowei Lou, Hye-young Paik, Wen Hu, Lina Yao arxiv

Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generativ…

DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech

2025-05-26 · Deok-Hyeon Cho, Hyung-Seok Oh, Seung-bin Kim, Seong-Whan Lee

Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods …

AttributeEmotional Speech SynthesisSpeech Synthesistext-to-speech+1

StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis

2024-09-24 · Zhiyong Chen, Xinnuo Li, Zhiqi Ai, Shugong Xu

We introduce StyleFusion-TTS, a prompt and/or audio referenced, style and speaker-controllable, zero-shot text-to-speech (TTS) synthesis system designed to enhance the editability and naturalness of current research lite…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis