paper-with-me

홈 › Papers

Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback

2026-01-27 · Siddhant Arora, Jinchuan Tian, Jiatong Shi, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, Shinji Watanabe arxiv

Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single semantic rewards applied at the utterance level. Such setups overlook the multi-dimensional and multi-modal nature of conversational quality, which encompasses semantic coherence, audio naturalness, speaker consistency, emotion alignment, and turn-taking behavior. Moreover, they are fundamentally mismatched with duplex spoken dialogue systems that generate responses incrementally, where agents must make decisions based on partial utterances. We address these limitations with the first multi-reward RLAIF framework for SDS, combining semantic, audio-quality, and emotion-consistency rewards. To align utterance-level preferences with incremental, blockwise decoding in duplex models, we apply turn-level preference sampling and aggregate per-block log-probabilities within a single DPO objective. We present the first systematic study of preference learning for improving SDS quality in both multi-turn Chain-of-Thought and blockwise duplex models, and release a multi-reward DPO dataset to support reproducible research. Experiments show that single-reward RLAIF selectively improves its targeted metric, while joint multi-reward training yields consistent gains across semantic quality and audio naturalness. These results highlight the importance of holistic, multi-reward alignment for practical conversational SDS.

📄 PDF Abstract BibTeX arXiv:2601.19063

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Self-supervised Dialogue Learning for Spoken Conversational Question Answering

2021-06-04 · Nuo Chen, Chenyu You, Yuexian Zou

In spoken conversational question answering (SCQA), the answer to the corresponding question is generated by retrieving and then analyzing a fixed spoken document, including multi-part conversations. Most SCQA systems ha…

Conversational Question Answeringcoreference-resolutionCoreference ResolutionQuestion Answering+2

Towards Data Distillation for End-to-end Spoken Conversational Question Answering

2020-10-18 · Chenyu You, Nuo Chen, Fenglin Liu, Dongchao Yang 외

In spoken question answering, QA systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via hum…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Conversational Question AnsweringQuestion Answering+2

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

2026-04-25 · Bhaskar Singh, Shobhit Banga, Mahima Manik, Pranav Sharma arxiv

Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first ope…

Text Generation

The Future of Spoken Dialogue Systems is in their Past: Long-Term Adaptive, Conversational Assistants

2012-06-01 · WS 2012 6 · David Schlangen
Language ModellingSpeech RecognitionSpeech SynthesisSpoken Dialogue Systems

End-to-end Spoken Conversational Question Answering: Task, Dataset and Model

2022-04-29 · Findings (NAACL) 2022 7 · Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge 외

In spoken question answering, the systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via hu…

4kConversational Question AnsweringQuestion AnsweringSpoken Language Understanding+1