paper-with-me

홈 › Papers

Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech

2025-09-18 · Taesoo Kim, Yongsik Jo, Hyunmin Song, Taehwan Kim arxiv

Human conversation involves language, speech, and visual cues, with each medium providing complementary information. For instance, speech conveys a vibe or tone not fully captured by text alone. While multimodal LLMs focus on generating text responses from diverse inputs, less attention has been paid to generating natural and engaging speech. We propose a human-like agent that generates speech responses based on conversation mood and responsive style information. To achieve this, we build a novel MultiSensory Conversation dataset focused on speech to enable agents to generate natural speech. We then propose a multimodal LLM-based model for generating text responses and voice descriptions, which are used to generate speech covering paralinguistic information. Experimental results demonstrate the effectiveness of utilizing both visual and audio modalities in conversation to generate engaging speech. The source code is available in https://github.com/kimtaesu24/MSenC

📄 PDF Abstract BibTeX arXiv:2509.14627

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents

2023-09-26 · Che-Jui Chang, Samuel S. Sohn, Sen Zhang, Rajath Jayashankar 외

Previous studies regarding the perception of emotions for embodied virtual agents have shown the effectiveness of using virtual characters in conveying emotions through interactions with humans. However, creating an auto…

Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents

2024-07-01 · Mehdi Arjmand, Farnaz Nouraei, Ian Steenstra, Timothy Bickmore

We introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the speak…

Emotional IntelligenceEmotion ClassificationHuman Interaction RecognitionLanguage Modelling+4

STICKERCONV: Generating Multimodal Empathetic Responses from Scratch

2024-01-20 · Yiqun Zhang, Fanheng Kong, Peidong Wang, Shuang Sun 외

Stickers, while widely recognized for enhancing empathetic communication in online interactions, remain underexplored in current empathetic dialogue research, notably due to the challenge of a lack of comprehensive datas…

2kEmpathetic Response GenerationResponse Generation

Situated and Interactive Multimodal Conversations

2020-06-02 · COLING 2020 8 · Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De 외

Next generation virtual assistants are envisioned to handle multimodal inputs (e.g., vision, memories of previous interactions, in addition to the user's utterances), and perform multimodal actions (e.g., displaying a ro…

Response Generation

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

2026-01-13 · Ashutosh Hathidara, Julien Yu, Vaishali Senthil, Sebastian Schreiber 외 arxiv

Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-user" prompting often yields verbose, unrea…