A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing
Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential of representation learning is yet to be realized, due to the challenging nature of generating high-quality speech. To address this problem, we propose our framework, Alignment-Aware Acoustic-Text Pretraining (A$^3$T), which reconstructs masked acoustic signals with text input and acoustic-text alignment during training. In this way, the pretrained model can generate high quality reconstructed spectrogram, which can be applied to the speech editing and unseen speaker TTS directly. Experiments show A$^3$T outperforms SOTA models on speech editing, and improves multi-speaker speech synthesis without the external speaker verification model.
Code (2)
Tasks
Representation LearningSpeaker Verificationspeech-recognitionSpeech RecognitionSpeech Representation LearningSpeech SynthesisSpeech-to-TextSpeech-to-Text TranslationTranslationSimilar Papers 제목 키워드 기반
EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG
Restoring speech communication from neural signals is a central goal of brain-computer interface research, yet EEG-based speech reconstruction remains challenging due to limited spatial resolution, susceptibility to nois…
Speech RecognitionTransfer LearningDomain AdaptationBridging What the Model Thinks and How It Speaks: Expressive Speech Generation via Self-Aware Intent-Realization Alignment
Speech Language Models (SLMs) exhibit strong semantic understanding, yet often fail to translate this capacity into expressive acoustic realization, producing speech with flattened prosody and misaligned emotion. We iden…
Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation
Whisper's robust performance in automatic speech recognition (ASR) is often attributed to its massive 680k-hour training set, an impractical scale for most researchers. In this work, we examine how linguistic and acousti…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDiversity+2ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment
Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains chall…
Audio GenerationJDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment
We propose Jointly trained Duration Informed Transformer (JDI-T), a feed-forward Transformer with a duration predictor jointly trained without explicit alignments in order to generate an acoustic feature sequence from an…
text-to-speechText to Speech