paper-with-me

홈 › Papers

Seeing and hearing what has not been said; A multimodal client behavior classifier in Motivational Interviewing with interpretable fusion

2023-09-25 · Lucie Galland, Catherine Pelachaud, Florian Pecune

Motivational Interviewing (MI) is an approach to therapy that emphasizes collaboration and encourages behavioral change. To evaluate the quality of an MI conversation, client utterances can be classified using the MISC code as either change talk, sustain talk, or follow/neutral talk. The proportion of change talk in a MI conversation is positively correlated with therapy outcomes, making accurate classification of client utterances essential. In this paper, we present a classifier that accurately distinguishes between the three MISC classes (change talk, sustain talk, and follow/neutral talk) leveraging multimodal features such as text, prosody, facial expressivity, and body expressivity. To train our model, we perform annotations on the publicly available AnnoMI dataset to collect multimodal information, including text, audio, facial expressivity, and body expressivity. Furthermore, we identify the most important modalities in the decision-making process, providing valuable insights into the interplay of different modalities during a MI conversation.

📄 PDF Abstract BibTeX arXiv:2309.14398

Code (1)

l-Galland/MALEFIC 공식 구현 pytorch

Tasks

Decision Making

Similar Papers 제목 키워드 기반

Seeing and Hearing Egocentric Actions: How Much Can We Learn?

2019-10-15 · Alejandro Cartas, Jordi Luque, Petia Radeva, Carlos Segura 외

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited nu…

Action Recognition

A Multimodal Simultaneous Interpretation Prototype: Who Said What

2022-09-01 · AMTA 2022 9 · Xiaolin Wang, Masao Utiyama, Eiichiro Sumita

“Who said what” is essential for users to understand video streams that have more than one speaker, but conventional simultaneous interpretation systems merely present “what was said” in the form of subtitles. Because th…

SentenceTAGTranslation

Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

2024-02-27 · CVPR 2024 1 · Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang 외

Video and audio content creation serves as the core technique for the movie industry and professional users. Recently, existing diffusion-based methods tackle video and audio generation separately, which hinders the tech…

Audio GenerationDenoising

M3-SLU: Evaluating Speaker-Attributed Reasoning in Multimodal Large Language Models

2025-10-22 · Yejin Kwon, Taewoo Kang, Hyunsoo Yoon, Changouk Kim arxiv

We present M3-SLU, a new multimodal large language model (MLLM) benchmark for evaluating multi-speaker, multi-turn spoken language understanding. While recent models show strong performance in speech and text comprehensi…

Spoken Language UnderstandingQuestion Answering

Responsibility in Extensive Form Games

2023-12-12 · Qi Shi

Two different forms of responsibility, counterfactual and seeing-to-it, have been extensively discussed in the philosophy and AI in the context of a single agent or multiple agents acting simultaneously. Although the gen…

counterfactualFormPhilosophy