paper-with-me

Papers

TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition

2023-05-23 · Hongfei Xue, Qijie Shao, Peikun Chen, Pengcheng Guo, Lei Xie, Jie Liu

UniSpeech has achieved superior performance in cross-lingual automatic speech recognition (ASR) by explicitly aligning latent representations to phoneme units using multi-task self-supervised learning. While the learned representations transfer well from high-resource to low-resource languages, predicting words directly from these phonetic representations in downstream ASR is challenging. In this paper, we propose TranUSR, a two-stage model comprising a pre-trained UniData2vec and a phoneme-to-word Transcoder. Different from UniSpeech, UniData2vec replaces the quantized discrete representations with continuous and contextual representations from a teacher model for phonetically-aware pre-training. Then, Transcoder learns to translate phonemes to words with the aid of extra texts, enabling direct word generation. Experiments on Common Voice show that UniData2vec reduces PER by 5.3% compared to UniSpeech, while Transcoder yields a 14.4% WER reduction compared to grapheme fine-tuning.

📄 PDF Abstract BibTeX arXiv:2305.13629

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learningspeech-recognitionSpeech RecognitionSpeech Representation Learning

Similar Papers 제목 키워드 기반

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

2026-06-30 · Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang 외 arxiv

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emot…

A unified front-end framework for English text-to-speech synthesis

2023-05-18 · Zelin Ying, Chen Li, Yu Dong, Qiuqiang Kong 외

The front-end is a critical component of English text-to-speech (TTS) systems, responsible for extracting linguistic features that are essential for a text-to-speech model to synthesize speech, such as prosodies and phon…

Speech SynthesisText Normalizationtext-to-speechText to Speech+1

Factorized RVQ-GAN For Disentangled Speech Tokenization

2025-06-18 · Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos 외

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge dist…

DisentanglementKnowledge DistillationSpeech Tokenization

LEARNING PHONEME-LEVEL DISCRETE SPEECH REPRESENTATION WITH WORD-LEVEL SUPERVISION

2021-09-29 · Liming Wang, Siyuan Feng, Mark A. Hasegawa-Johnson, Chang D. Yoo

Phonemes are defined by their relationship to words: changing a phoneme changes the word. Learning a phoneme inventory with little supervision has been a long-standing challenge with important applications to under-reso…

Representation LearningSelf-Supervised Learning

Comparing phonemes and visemes with DNN-based lipreading

2018-05-08 · Kwanchiva Thangthai, Helen L. Bear, Richard Harvey

There is debate if phoneme or viseme units are the most effective for a lipreading system. Some studies use phoneme units even though phonemes describe unique short sounds; other studies tried to improve lipreading accur…

DecoderLipreading