paper-with-me

Papers

A Dynamic Speaker Model for Conversational Interactions

2019-06-01 · NAACL 2019 6 · Hao Cheng, Hao Fang, Mari Ostendorf

Individual differences in speakers are reflected in their language use as well as in their interests and opinions. Characterizing these differences can be useful in human-computer interaction, as well as analysis of human-human conversations. In this work, we introduce a neural model for learning a dynamically updated speaker embedding in a conversational context. Initial model training is unsupervised, using context-sensitive language generation as an objective, with the context being the conversation history. Further fine-tuning can leverage task-dependent supervised training. The learned neural representation of speakers is shown to be useful for content ranking in a socialbot and dialog act prediction in human-human conversations.

📄 PDF Abstract BibTeX

Code (1)

hao-cheng/dynamic_speaker_model 공식 구현 tf

Tasks

modelText Generation

Similar Papers 제목 키워드 기반

Speaker-Aware Simulation Improves Conversational Speech Recognition

2026-02-04 · Máté Gedeon, Péter Mihajlik arxiv

Automatic speech recognition (ASR) for conversational speech remains challenging due to the limited availability of large-scale, well-annotated multi-speaker dialogue data and the complex temporal dynamics of natural int…

Dialogue GenerationSpeech RecognitionData Augmentation

ESIHGNN: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition

2024-05-07 · Xupeng Zha, Huan Zhao, Zixing Zhang

Conversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an ``event'') during a conversation. Existing graph-based methods mainly focus on event interactions to compr…

Emotion RecognitionGraph Neural Network

Multi-human Interactive Talking Dataset

2025-08-05 · Zeyu Zhu, Weijia Wu, Mike Zheng Shou arxiv

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, w…

Video Generation

Speaker-Aware Discourse Parsing on Multi-Party Dialogues

2022-10-01 · COLING 2022 10 · Nan Yu, Guohong Fu, Min Zhang

Discourse parsing on multi-party dialogues is an important but difficult task in dialogue systems and conversational analysis. It is believed that speaker interactions are helpful for this task. However, most previous re…

Discourse Parsing

deep learning of segment-level feature representation for speech emotion recognition in conversations

2023-02-05 · Jiachen Luo, Huy Phan, Joshua Reiss

Accurately detecting emotions in conversation is a necessary yet challenging task due to the complexity of emotions and dynamics in dialogues. The emotional state of a speaker can be influenced by many different factors,…

Emotion RecognitionSpeech Emotion Recognition