paper-with-me

홈 › Papers

WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

2026-04-16 · Yifu Chen, Shengpeng Ji, Qian Chen, Tianle Liang, Yangzhuo Li, Ziqing Wang, Wen Wang, Jingyu Lu, Haoxiao Wang, Xueyi Pu, Fan Zhuo, Zhou Zhao arxiv

End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems. However, the intelligence and expressiveness of current open-source spoken dialogue models often remain below expectations. Motivated by the success of online reinforcement learning(RL) in other domains, one might attempt to directly apply preference optimization to spoken dialogue models, yet this transfer is non-trivial. We analyze these obstacles from the perspectives of reward modeling and rollout sampling, focusing on how sparse preference supervision interacts with dense speech generation under shared-parameter updates. Based on the analysis, we propose a modality-aware adaptive post-training recipe that makes RL practical for spoken dialogue: it constrains preference updates to the semantic channel and improves acoustic behavior via explicit anchoring, while dynamically regulating their mixture from rollout statistics to avoid unreliable preference gradients. We evaluate the method across multiple spoken dialogue benchmarks and representative architectures, and observe consistent improvements in semantic quality and speech expressiveness.

📄 PDF Abstract BibTeX arXiv:2604.14932

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

2026-09-15 · Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu 외 arxiv

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can s…

WavChat: A Survey of Spoken Dialogue Models

2024-11-15 · Shengpeng Ji, Yifu Chen, Minghui Fang, Jialong Zuo 외

Recent advancements in spoken dialogue models, exemplified by systems like GPT-4o, have captured significant attention in the speech domain. Compared to traditional three-tier cascaded spoken dialogue models that compris…

speech-recognitionSpeech RecognitionSpoken Dialogue SystemsSurvey+2

FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs

2026-04-20 · Yun Hong, Yan Zhou, Yang Feng arxiv

Empathy is essential for fostering natural interactions in spoken dialogue systems, as it enables machines to recognize the emotional tone of human speech and deliver empathetic responses. Recent research has made signif…

Speech Emotion Recognition

Emotional Speech Corpus for Persuasive Dialogue System

2020-05-01 · LREC 2020 5 · Sara Asai, Koichiro Yoshino, Seitaro Shinagawa, Sakriani Sakti 외

Expressing emotion is known as an efficient way to persuade one{'}s dialogue partner to accept one{'}s claim or proposal. Emotional expression in speech can express the speaker{'}s emotion more directly than using only e…

End-to-End Evaluation of a Spoken Dialogue System for Learning Basic Mathematics

2022-11-07 · Eda Okur, Saurav Sahay, Roddy Fuentes Alba, Lama Nachman

The advances in language-based Artificial Intelligence (AI) technologies applied to build educational applications can present AI for social-good opportunities with a broader positive impact. Across many disciplines, enh…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Intent RecognitionMath+3