paper-with-me

Papers

Improving Speech Representation Learning via Speech-level and Phoneme-level Masking Approach

2022-10-25 · xulong Zhang, Jianzong Wang, Ning Cheng, Kexin Zhu, Jing Xiao

Recovering the masked speech frames is widely applied in speech representation learning. However, most of these models use random masking in the pre-training. In this work, we proposed two kinds of masking approaches: (1) speech-level masking, making the model to mask more speech segments than silence segments, (2) phoneme-level masking, forcing the model to mask the whole frames of the phoneme, instead of phoneme pieces. We pre-trained the model via these two approaches, and evaluated on two downstream tasks, phoneme classification and speaker recognition. The experiments demonstrated that the proposed masking approaches are beneficial to improve the performance of speech representation.

📄 PDF Abstract BibTeX arXiv:2210.13805

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningSpeaker RecognitionSpeech Representation Learning

Similar Papers 제목 키워드 기반

Exploring Phoneme-Level Speech Representations for End-to-End Speech Translation

2019-06-04 · ACL 2019 7 · Elizabeth Salesky, Matthias Sperber, Alan W. black

Previous work on end-to-end translation from speech has primarily used frame-level features as speech representations, which creates longer, sparser sequences than text. We show that a naive method to create compressed p…

Translation

Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes

2024-12-17 · Kuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang 외

Recent advancements in text-to-speech and speech conversion technologies have enabled the creation of highly convincing synthetic speech. While these innovations offer numerous practical benefits, they also cause signifi…

DeepFake DetectionFace SwappingGraph Attentiontext-to-speech+1

LEARNING PHONEME-LEVEL DISCRETE SPEECH REPRESENTATION WITH WORD-LEVEL SUPERVISION

2021-09-29 · Liming Wang, Siyuan Feng, Mark A. Hasegawa-Johnson, Chang D. Yoo

Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a long-standing challenge with important applications to under-reso…

Representation LearningSelf-Supervised Learning

DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition

2025-01-31 · Wonjun Lee, Solee Im, Heejin Do, Yunsu Kim 외

Dysarthric speech recognition often suffers from performance degradation due to the intrinsic diversity of dysarthric severity and extrinsic disparity from normal speech. To bridge these gaps, we propose a Dynamic Phonem…

Contrastive LearningDiversityspeech-recognitionSpeech Recognition

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

2026-06-30 · Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang 외 arxiv

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emot…