paper-with-me

Speech Synthesis

5개 벤치마크 · 논문 1,383편 · 이 태스크의 논문 보기 →

Benchmarks

LibriTTS

결과 45개

North American English

결과 21개

LJSpeech

결과 12개

Mandarin Chinese

결과 9개

Most implemented

WaveNet: A Generative Model for Raw Audio

2016-09-12 · 구현 62개

Papers

What Did I Just Say? Self-Listening for Full-Duplex Speech Models

2026-09-04 · Xuanning Zhou, Junyi Ao, Xiaotong Liu, Tom Ko 외 hf

Full-duplex spoken language models can listen and speak simultaneously, enabling them to handle interruptions and backchannels in human conversation. However, text generation, speech synthesis, and audio playback proceed…

Speech SynthesisText Generation

SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

2026-08-18 · Shicheng Ma, Wenqian Cui, Irwin King arxiv

Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding t…

Sentiment AnalysisSpeech RecognitionSpeech Synthesis

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

2026-08-14 · Xueqi Wang, Zhigang Wang, Runqing Zhang, Zhenqi Jia 외 arxiv

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level ret…

Emotion Recognition in ConversationContrastive LearningSpeech Synthesis

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

2026-08-13 · Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor, Shehzeen Hussain 외 arxiv

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-spe…

Speech Synthesis

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

2026-08-09 · Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang 외 arxiv

Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency respons…

Speech Synthesis

Teffic-Audio: Tell Fact from Fiction

2026-07-30 · Wan Lin, Li Wang, Jindong Wang, Kunyu Feng 외 arxiv

Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis. The resulting spoof…

DeepFake DetectionSpeech SynthesisVoice Conversion

전체 1,383편 보기 →