paper-with-me

Papers

Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

2024-02-20 · Guan-Ting Lin, Cheng-Han Chiang, Hung-Yi Lee

In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing paralinguistic and prosodic information, mark the most significant difference between text and speech modality. When using text-only LLMs to model spoken dialogue, text-only LLMs cannot give different responses based on the speaking style of the current turn. In this paper, we focus on enabling LLMs to listen to the speaking styles and respond properly. Our goal is to teach the LLM that "even if the sentences are identical if they are spoken in different styles, their corresponding responses might be different". Since there is no suitable dataset for achieving this goal, we collect a speech-to-speech dataset, StyleTalk, with the following desired characteristics: when two current speeches have the same content but are spoken in different styles, their responses will be different. To teach LLMs to understand and respond properly to the speaking styles, we propose the Spoken-LLM framework that can model the linguistic content and the speaking styles. We train Spoken-LLM using the StyleTalk dataset and devise a two-stage training pipeline to help the Spoken-LLM better learn the speaking styles. Based on extensive experiments, we show that Spoken-LLM outperforms text-only baselines and prior speech LLMs methods.

📄 PDF Abstract BibTeX arXiv:2402.12786

Code (1)

daniellin94144/styletalk 공식 구현

Tasks

Sentence

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Advancing Multi-Party Dialogue Framework with Speaker-ware Contrastive Learning

2025-01-20 · Zhongtian Hu, Qi He, Ronghan Li, Meng Zhao 외

Multi-party dialogues, common in collaborative scenarios like brainstorming sessions and negotiations, pose significant challenges due to their complexity and diverse speaker roles. Current methods often use graph neural…

Contrastive LearningDialogue GenerationResponse Generation

SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning

2024-08-25 · Chien-yu Huang, Min-Han Shih, Ke-Han Lu, Chi-Yuan Hsiao 외

Instruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirabl…

Emotion Recognition

Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation

2024-08-18 · Xukun Zhou, Fengxin Li, Ziqiao Peng, Kejian Wu 외

Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with …

3D Face AnimationMeta-LearningModel Optimization

Advancing Automated Speaking Assessment Leveraging Multifaceted Relevance and Grammar Information

2025-06-19 · Hao-Chien Lu, Jhen-Ke Lin, Hong-Yun Lin, Chung-Chun Wang 외

Current automated speaking assessment (ASA) systems for use in multi-aspect evaluations often fail to make full use of content relevance, overlooking image or exemplar cues, and employ superficial grammar analysis that l…

Learning to mirror speaking styles incrementally

2020-03-05 · Siyi Liu, Ziang Leng, Derry Wijaya

Mirroring is the behavior in which one person subconsciously imitates the gesture, speech pattern, or attitude of another. In conversations, mirroring often signals the speakers enjoyment and engagement in their communic…