paper-with-me

Text to Speech

2개 벤치마크 · 논문 1,433편 · 이 태스크의 논문 보기 →

Benchmarks

.

결과 2개

^(#$!@#$)(()))******

결과 2개

Most implemented

WaveNet: A Generative Model for Raw Audio

2016-09-12 · 구현 62개

Efficient Neural Audio Synthesis

2018-02-23 · 구현 16개

Papers

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

2026-05-21 · Luca Modica, Filip Landin, Mehrdad Farahani, Livia Qian 외 arxiv

In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges about how model-internal mechanisms are similar and different when oper…

Text to Speech

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

2026-04-23 · Srija Anand, Ashwin Sankar, Ishvinder Sethi, Aaditya Pareek 외 arxiv

Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introduces high variance due to linguistic diversity and multidimensional n…

Text to Speech

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

2026-02-17 · Ahmed Khaled Khamis, Hesham Ali arxiv

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …

Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to Speech

ManchuTTS: Towards High-Quality Manchu Speech Synthesis via Flow Matching and Hierarchical Text Representation

2025-12-27 · Suhua Wang, Zifan Wang, Xiaoxin Sun, D. J. Wang 외 arxiv

As an endangered language, Manchu presents unique challenges for speech synthesis, including severe data scarcity and strong phonological agglutination. This paper proposes ManchuTTS(Manchu Text to Speech), a novel appro…

Data AugmentationSpeech SynthesisText to Speech

Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization

2025-09-29 · Jiacheng Shi, Hongfei Du, Yangfan He, Y. Alicia Hong 외 arxiv

Emotional text-to-speech seeks to convey affect while preserving intelligibility and prosody, yet existing methods rely on coarse labels or proxy classifiers and receive only utterance-level feedback. We introduce Emotio…

Text to Speech

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

2025-09-25 · Sitong Cheng, Weizhen Bian, Xinsheng Wang, Ruibin Yuan 외 arxiv

The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered…

Speech-to-Speech TranslationText to Speech

전체 1,433편 보기 →