paper-with-me

Papers

PAMA-TTS: Progression-Aware Monotonic Attention for Stable Seq2Seq TTS With Accurate Phoneme Duration Control

2021-10-09 · Yunchao He, Jian Luan, Yujun Wang

Sequence expansion between encoder and decoder is a critical challenge in sequence-to-sequence TTS. Attention-based methods achieve great naturalness but suffer from unstable issues like missing and repeating phonemes, not to mention accurate duration control. Duration-informed methods, on the contrary, seem to easily adjust phoneme duration but show obvious degradation in speech naturalness. This paper proposes PAMA-TTS to address the problem. It takes the advantage of both flexible attention and explicit duration models. Based on the monotonic attention mechanism, PAMA-TTS also leverages token duration and relative position of a frame, especially countdown information, i.e. in how many future frames the present phoneme will end. They help the attention to move forward along the token sequence in a soft but reliable control. Experimental results prove that PAMA-TTS achieves the highest naturalness, while has on-par or even better duration controllability than the duration-informed model.

📄 PDF Abstract BibTeX arXiv:2110.04486

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

Consistent Style Transfer

2022-01-06 · Xuan Luo, Zhen Han, Lingkang Yang, Lingling Zhang

Recently, attentional arbitrary style transfer methods have been proposed to achieve fine-grained results, which manipulates the point-wise similarity between content and style features for stylization. However, the atte…

Style Transfer

Pan-cancer Histopathology WSI Pre-training with Position-aware Masked Autoencoder

2024-07-10 · Kun Wu, Zhiguo Jiang, Kunming Tang, Jun Shi 외

Large-scale pre-training models have promoted the development of histopathology image analysis. However, existing self-supervised methods for histopathology images primarily focus on learning patch features, while there …

Cancer ClassificationPositionRepresentation LearningSelf-Supervised Learning

PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies

2026-06-25 · Jiayu Yang, Tao Yang, Xiang Chang, Fei Chao 외 arxiv

Reliable action generation for multi-stage robotic manipulation remains challenging for Vision-Language-Action (VLA) models. While existing flow-matching VLA policies offer strong multimodal grounding and generalization,…

StableEmit: Selection Probability Discount for Reducing Emission Latency of Streaming Monotonic Attention ASR

2021-07-01 · Hirofumi Inaguma, Tatsuya Kawahara

While attention-based encoder-decoder (AED) models have been successfully extended to the online variants for streaming automatic speech recognition (ASR), such as monotonic chunkwise attention (MoChA), the models still …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Boundary DetectionDecoder+2

Efficient Monotonic Multihead Attention

2023-12-07 · Xutai Ma, Anna Sun, Siqi Ouyang, Hirofumi Inaguma 외

We introduce the Efficient Monotonic Multihead Attention (EMMA), a state-of-the-art simultaneous translation model with numerically-stable and unbiased monotonic alignment estimation. In addition, we present improved tra…

Simultaneous Speech-to-Text TranslationSpeech-to-TextSpeech-to-Text TranslationTranslation