paper-with-me

Papers

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

2026-01-14 · Rajarshi Roy, Jonathan Raiman, Sang-gil Lee, Teodor-Dumitru Ene, Robert Kirby, Sungwon Kim, Jaehyeon Kim, Bryan Catanzaro arxiv

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, role-driven real-world applications and personalized interactions. In this work, we introduce PersonaPlex, a duplex conversational speech model that incorporates hybrid system prompts, combining role conditioning with text prompts and voice cloning with speech samples. PersonaPlex is trained on a large-scale synthetic dataset of paired prompts and user-agent conversations, generated with open-source large language models (LLM) and text-to-speech (TTS) models. To evaluate role conditioning in real-world settings, we extend the Full-Duplex-Bench benchmark beyond a single assistant role to multi-role customer service scenarios. Experiments show that PersonaPlex achieves strong role-conditioned behavior, voice-conditioned speech, and natural conversational responsiveness, surpassing state-of-the-art duplex speech models and hybrid large language model-based speech systems in role adherence, speaker similarity, latency, and naturalness.

📄 PDF Abstract BibTeX arXiv:2602.06053

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EchoChain: A Full-Duplex Benchmark for State-Update Reasoning Under Interruptions

2026-04-08 · Smit Nautambhai Modi, Gandharv Mahajan, Marc Wetter, Randall Welles arxiv

Real-time voice assistants must revise task state when users interrupt mid-response, but existing spoken-dialog benchmarks largely evaluate turn-based interaction and miss this failure mode. We introduce EchoChain, a con…

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

2026-06-09 · Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez, Eugene Kharitonov arxiv

Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-lev…

Reinforcement LearningDialogue Evaluation

DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization

2026-03-10 · Jianing Yang, Yusuke Fujita, Yui Sudo arxiv

Spoken dialog systems with cascaded ASR-LLM-TTS modules retain strong LLM intelligence, but VAD segmentation often forces half-duplex turns and brittle control. On the other hand, VAD-free end-to-end model support full-d…

$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

2026-03-14 · Soham Ray, Keshav Dhandhania, Victor Barres, Karthik Narasimhan arxiv

Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We i…

Raon-Speech Technical Report

2026-04-08 · Beomsoo Kim, Changho Choi, Dohyun Kim, Dongki Lee 외 arxiv

We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat, a high-performing full-duplex extension …

Knowledge DistillationQuestion Answering