paper-with-me

Papers

Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis

2026-04-24 · Xiang Zhang, Xiaotian Li, Taoyue Wang, Nan Bi, Xin Zhou, Cody Zhou, Zoie Wang, Andrew Yang, Yuming Su, Jeff Cohn, Qiang Ji, Lijun Yin arxiv

Social interactions dominate our perceptions of the world and shape our daily behavior by attaching social meaning to acts as simple and spontaneous as gestures, facial expressions, voice, and speech. People mimic and otherwise respond to each other's postures, facial expressions, mannerisms, and other verbal and nonverbal behavior, and form appraisals or evaluations in the process. Yet, no publicly-available dataset includes multimodal recordings and self-report measures of multiple persons in social interaction. Dyadic recordings and annotation are lacking. We present a new data corpus of multimodal dyadic interaction (45 dyads, 90 persons) that includes synchronized multi-modality behavior (2D face video, 3D face geometry, thermal spectrum dynamics, voice and speech behavior, physiology (PPG, EDA, heart-rate, blood pressure, and respiration), and self-reported affect of all participants in a communicative interaction scenario. Two types of dyads are included: persons with shared past history and strangers. Annotations include social signals, agreement, disagreement, and neutral stance. With a potent emotion induction, these multimodal data will enable novel modeling of multimodal interpersonal behavior. We present extensive experiments to evaluate multimodal dyadic communication of dyads with and without interpersonal history, and their affect. This new database will make multimodal modeling of social interaction never possible before. The dataset includes 20TB of multimodal data to share with the research community.

📄 PDF Abstract BibTeX arXiv:2604.22739

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech

2026-03-03 · Thanushi Withanage, Elizabeth Redcay, Carol Espy-Wilson arxiv

Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, …

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

2026-08-22 · Ahmet Tuğrul Bayrak, Fatma Nur Korkmaz, Bekir Berker Türker, Mustafa Sertaç Türkel 외 hf

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language mod…

Binary Classification

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

2025-05-27 · Cheng Luo, Jianghui Wang, Bing Li, Siyang Song 외

In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task that aims to online generate synchronized verbal and non-verbal listener feedback, conditioned on the speaker's multi…

Audio-Visual SynchronizationConversational Response GenerationLarge Language ModelMultimodal Large Language Model+1

Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos

2018-06-01 · NAACL 2018 6 · Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria 외

Emotion recognition in conversations is crucial for the development of empathetic machines. Present methods mostly ignore the role of inter-speaker dependency relations while classifying emotions in conversations. In thi…

Emotion RecognitionEmotion Recognition in Conversation

Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition

2026-09-09 · Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte arxiv

Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the complexity of natural conversation, which involves overlapping speech, …

Audio-Visual Speech Recognition