paper-with-me

홈 › Papers

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

2026-05-07 · Jiacheng Xu, Heting Gao, Liufei Xie, Zhenchuan Yang, Lijiang Li, Yiting Chen, Bin Zhang, Meng Chen, Chaoyu Fu, Weifeng Zhao, Wenjiang Zhou arxiv

Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humming a song, which we formalize as role-playing and singing. We present VITA-QinYu, the first expressive end-to-end (E2E) spoken language model (SLM) that goes beyond natural conversation to support both role-playing and singing generation. VITA-QinYu adopts a hybrid speech-text paradigm that extends interleaved text-audio modeling with multi-codebook audio tokens, a design enabling richer paralinguistic representation while preserving a clear separation between modalities to avoid interference. We further develop a comprehensive data generation pipeline to synthesize a total of 15.8K hours of natural conversation, role-playing, and singing data for training. VITA-QinYu demonstrates superior expressiveness, outperforming peer SLMs by 7 percentage points on objective role-playing benchmarks, and surpassing peer models by 0.13 points on a 5-point MOS scale for singing. Simultaneously, it achieves state-of-the-art conversational accuracy and fluency, exceeding prior SLMs by 1.38 and 4.98 percentage points on the C3 and URO benchmarks, respectively. We open-source our code and models and provide an easy-to-use demo with full-stack support for streaming and full-duplex interaction.

📄 PDF Abstract BibTeX arXiv:2605.06765

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

2026-09-15 · Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu 외 arxiv

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can s…

Re:Member: Emotional Question Generation from Personal Memories

2025-10-21 · Zackary Rackauckas, Nobuaki Minematsu, Julia Hirschberg arxiv

We present Re:Member, a system that explores how emotionally expressive, memory-grounded interaction can support more engaging second language (L2) learning. By drawing on users' personal videos and generating stylized s…

Question Generation

Celtic CALL: strengthening the vital role of education for language transmission

2022-06-01 · CLTW (LREC) 2022 6 · Neasa Ní Chiaráin, Madeleine Comtois, Oisín Nolan, Neimhin Robinson-Gunning 외

In this paper, we present the Irish language learning platform, An Sc ́eala ́ı, an intelligent Computer-Assisted Language Learning (iCALL) system which incorporates speech and language technologies in ways that promote t…

PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text

2021-04-01 · EACL 2021 2 · Varnith Chordia

Text transcripts without punctuation or sentence boundaries are hard to comprehend for both humans and machines. Punctuation marks play a vital role by providing meaning to the sentence and incorrect use or placement of …

Language ModellingPunctuation RestorationSentenceTranslation

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

2026-04-16 · Yifu Chen, Shengpeng Ji, Qian Chen, Tianle Liang 외 arxiv

End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems. However, the intelligence and expressive…

Reinforcement Learning