paper-with-me

Papers

Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration

2024-12-11 · Haowei Lou, Helen Paik, Wen Hu, Lina Yao

Recent advancements in text-to-speech (TTS) systems, such as FastSpeech and StyleSpeech, have significantly improved speech generation quality. However, these models often rely on duration generated by external tools like the Montreal Forced Aligner, which can be time-consuming and lack flexibility. The importance of accurate duration is often underestimated, despite their crucial role in achieving natural prosody and intelligibility. To address these limitations, we propose a novel Aligner-Guided Training Paradigm that prioritizes accurate duration labelling by training an aligner before the TTS model. This approach reduces dependence on external tools and enhances alignment accuracy. We further explore the impact of different acoustic features, including Mel-Spectrograms, MFCCs, and latent features, on TTS model performance. Our experimental results show that aligner-guided duration labelling can achieve up to a 16\% improvement in word error rate and significantly enhance phoneme and tone alignment. These findings highlight the effectiveness of our approach in optimizing TTS systems for more natural and intelligible speech generation.

📄 PDF Abstract BibTeX arXiv:2412.08112

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

2026-09-09 · Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao 외 arxiv

Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical variation and subtle d…

Tem-adapter: Adapting Image-Text Pretraining for Video Question Answer

2023-08-16 · ICCV 2023 1 · Guangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang 외

Video-language pre-trained models have shown remarkable success in guiding video question-answering (VideoQA) tasks. However, due to the length of video sequences, training large-scale video-based models incurs considera…

DecoderQuestion AnsweringVideo Question Answering

Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction

2025-01-09 · Hantao Lou, Jiaming Ji, Kaile Wang, Yaodong Yang

The rapid advancement of large language models (LLMs) has led to significant improvements in their capabilities, but also to increased concerns about their alignment with human values and intentions. Current alignment st…

MathSentence

NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

2024-05-02 · Gerald Shen, Zhilin Wang, Olivier Delalleau, Jiaqi Zeng 외

Aligning Large Language Models (LLMs) with human values and preferences is essential for making them helpful and safe. However, building efficient tools to perform alignment can be challenging, especially for the largest…

modelparameter-efficient fine-tuning

MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models

2024-03-25 · Kailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 외

Recent advancements in large language models (LLMs) focus on aligning to heterogeneous human expectations and values via multi-objective preference alignment. However, existing methods are dependent on the policy model p…

GPUIn-Context Learning