paper-with-me

Papers

A$^3$T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing

2022-03-18 · He Bai, Renjie Zheng, Junkun Chen, Xintong Li, Mingbo Ma, Liang Huang

Recently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation. However, all the above tasks are in the direction of speech understanding, but for the inverse direction, speech synthesis, the potential of representation learning is yet to be realized, due to the challenging nature of generating high-quality speech. To address this problem, we propose our framework, Alignment-Aware Acoustic-Text Pretraining (A$^3$T), which reconstructs masked acoustic signals with text input and acoustic-text alignment during training. In this way, the pretrained model can generate high quality reconstructed spectrogram, which can be applied to the speech editing and unseen speaker TTS directly. Experiments show A$^3$T outperforms SOTA models on speech editing, and improves multi-speaker speech synthesis without the external speaker verification model.

📄 PDF Abstract BibTeX arXiv:2203.09690

Code (2)

PaddlePaddle/PaddleSpeech/tree/develop/examples/vctk/ernie_sat paddle
richardbaihe/a3t pytorch

Tasks

Representation LearningSpeaker Verificationspeech-recognitionSpeech RecognitionSpeech Representation LearningSpeech SynthesisSpeech-to-TextSpeech-to-Text TranslationTranslation

Similar Papers 제목 키워드 기반

EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG

2025-12-14 · Hanbeot Park, Yunjeong Cho, Hunhee Kim arxiv

Restoring speech communication from neural signals is a central goal of brain-computer interface research, yet EEG-based speech reconstruction remains challenging due to limited spatial resolution, susceptibility to nois…

Speech RecognitionTransfer LearningDomain Adaptation

Bridging What the Model Thinks and How It Speaks: Expressive Speech Generation via Self-Aware Intent-Realization Alignment

2026-04-13 · Kuang Wang, Lai Wei, Ping Lin, Qibing Bai 외 arxiv

Speech Language Models (SLMs) exhibit strong semantic understanding, yet often fail to translate this capacity into expressive acoustic realization, producing speech with flattened prosody and misaligned emotion. We iden…

Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation

2025-05-27 · Dancheng Liu, Amir Nassereldine, Chenhui Xu, JinJun Xiong

Whisper's robust performance in automatic speech recognition (ASR) is often attributed to its massive 680k-hour training set, an impractical scale for most researchers. In this work, we examine how linguistic and acousti…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationDiversity+2

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

2026-05-29 · Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee arxiv

Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains chall…

Audio Generation

JDI-T: Jointly trained Duration Informed Transformer for Text-To-Speech without Explicit Alignment

2020-05-15 · Dan Lim, Won Jang, Gyeonghwan O, Heayoung Park 외

We propose Jointly trained Duration Informed Transformer (JDI-T), a feed-forward Transformer with a duration predictor jointly trained without explicit alignments in order to generate an acoustic feature sequence from an…

text-to-speechText to Speech