paper-with-me

Papers

Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?

2024-10-25 · Opeyemi Osakuade, Simon King

Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models, are widely used, especially where there are limited data for the downstream task, such as for a low-resource language. Typically, discretization of speech into a sequence of symbols is achieved by unsupervised clustering of the latents from an SSL model. Our study evaluates whether discrete symbols - found using k-means - adequately capture tone in two example languages, Mandarin and Yoruba. We compare latent vectors with discrete symbols, obtained from HuBERT base, MandarinHuBERT, or XLS-R, for vowel and tone classification. We find that using discrete symbols leads to a substantial loss of tone information, even for language-specialised SSL models. We suggest that discretization needs to be task-aware, particularly for tone-dependent downstream tasks.

📄 PDF Abstract BibTeX arXiv:2410.19935

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

2019-10-12 · ICLR 2020 1 · Alexei Baevski, Steffen Schneider, Michael Auli

We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quanti…

ClusteringGeneral ClassificationSelf-Supervised Learningspeech-recognition+1

A Comparison of Discrete and Soft Speech Units for Improved Voice Conversion

2021-11-03 · Benjamin van Niekerk, Marc-André Carbonneau, Julian Zaïdi, Mathew Baas 외

The goal of voice conversion is to transform source speech into a target voice, keeping the content unchanged. In this paper, we focus on self-supervised representation learning for voice conversion. Specifically, we com…

Representation LearningVoice Conversion

Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer

2023-09-14 · Yongqi Wang, Jionghao Bai, Rongjie Huang, RuiQi Li 외

Direct speech-to-speech translation (S2ST) with discrete self-supervised representations has achieved remarkable accuracy, but is unable to preserve the speaker timbre of the source speech. Meanwhile, the scarcity of hig…

In-Context LearningLanguage ModelingLanguage ModellingSpeech-to-Speech Translation+2

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model

2024-12-04 · Joonyong Park, Daisuke Saito, Nobuaki Minematsu

We examine the text-free speech representations of raw audio obtained from a self-supervised learning (SSL) model by analyzing the synthesized speech using the SSL representations instead of conventional text representat…

Self-Supervised LearningSpeech Synthesis

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

2021-04-01 · Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov 외

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representations for speech content, prosodic inform…

DisentanglementRepresentation LearningResynthesisSpeaker Identification+1