paper-with-me

홈 › Papers

RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis

2025-10-16 · Qing Yang, Zhenghao Liu, Yangfan Du, Pengcheng Huang, Tong Xiao arxiv

Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate objectives that fail to adequately capture perceptual emotional quality. As a result, the generated speech, while semantically accurate, often lacks expressive and emotionally rich characteristics. To address these limitations, we propose RLAIF-SPA, a novel framework that integrates Reinforcement Learning from AI Feedback (RLAIF) to directly optimize both emotional expressiveness and intelligibility without human supervision. Specifically, RLAIF-SPA incorporates Automatic Speech Recognition (ASR) to provide semantic accuracy feedback, while leveraging structured reward modeling to evaluate prosodic-emotional consistency. RLAIF-SPA enables more precise and nuanced control over expressive speech generation along four structured evaluation dimensions: Structure, Emotion, Speed, and Tone. Extensive experiments on Libri-Speech, MELD, and Mandarin ESD datasets demonstrate consistent gains across clean read speech, conversational dialogue, and emotional speech. On Libri-Speech, RLAIF-SPA consistently outperforms Chat-TTS, achieving a 26.1% reduction in word error rate, a 9.1% improvement in SIM-O, and over 10% gains in human subjective evaluations.

📄 PDF Abstract BibTeX arXiv:2510.14628

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSpeech RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

2024-05-27 · CVPR 2025 1 · Tianyu Yu, Haoye Zhang, Qiming Li, Qixin Xu 외

Traditional feedback learning for hallucination reduction relies on labor-intensive manual labeling or expensive proprietary models. This leaves the community without foundational knowledge about how to build high-qualit…

HallucinationImage CaptioningObject HallucinationVisual Question Answering

Multi-objective Reinforcement learning from AI Feedback

2024-06-11 · Marcus Williams

This paper presents Multi-Objective Reinforcement Learning from AI Feedback (MORLAIF), a novel approach to improving the alignment and performance of language models trained using reinforcement learning from AI feedback …

Language ModelingLanguage ModellingMulti-Objective Reinforcement Learningreinforcement-learning+1

Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback

2026-01-27 · Siddhant Arora, Jinchuan Tian, Jiatong Shi, Hayato Futami 외 arxiv

Reinforcement learning from human or AI feedback (RLHF/RLAIF) for speech-in/speech-out dialogue systems (SDS) remains underexplored, with prior work largely limited to single semantic rewards applied at the utterance lev…

Reinforcement Learning

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback

2025-10-02 · Derek Shi, Ruben Glatt, Christine Klymko, Shubham Mohole 외 arxiv

Recent advances in large video-language models (VLMs) rely on extensive fine-tuning techniques that strengthen alignment between textual and visual comprehension. Leading pipelines typically pair supervised fine-tuning (…

Reinforcement Learning

Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

2024-02-06 · Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang 외

Recent advancements in large language models have influenced the development of video large multimodal models (VLMMs). The previous approaches for VLMMs involved Supervised Fine-Tuning (SFT) with instruction-tuned datase…

Video-based Generative Performance Benchmarking