paper-with-me

홈 › Papers

Deep Learning Based Multi-modal Addressee Recognition in Visual Scenes with Utterances

2018-09-12 · Thao Minh Le, Nobuyuki Shimizu, Takashi Miyazaki, Koichi Shinoda

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including those outdoors, in cafeterias, and hospitals. Because previous studies typically focused only on pre-specified tasks with limited conversational situations such as controlling smart homes, we created a mock dataset called Addressee Recognition in Visual Scenes with Utterances (ARVSU) that contains a vast body of image variations in visual scenes with an annotated utterance and a corresponding addressee for each scenario. We also propose a multi-modal deep-learning-based model that takes different human cues, specifically eye gazes and transcripts of an utterance corpus, into account to predict the conversational addressee from a specific speaker's view in various real-life conversational scenarios. To the best of our knowledge, we are the first to introduce an end-to-end deep learning model that combines vision and transcripts of utterance for addressee recognition. As a result, our study suggests that future addressee recognition can reach the ability to understand human intention in many social situations previously unexplored, and our modality dataset is a first step in promoting research in this field.

📄 PDF Abstract BibTeX arXiv:1809.04288

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue

2025-01-28 · Koji Inoue, Divesh Lala, Mikey Elmers, Keiko Ochi 외

Handling multi-party dialogues represents a significant step for advancing spoken dialogue systems, necessitating the development of tasks specific to multi-party interactions. To address this challenge, we are construct…

Language ModelingLanguage ModellingLarge Language ModelSpoken Dialogue Systems

Evaluating Large Language Models Abilities for Addressee, Turn-change, and Next Speaker Prediction in Meetings

2026-06-16 · Ryo Fukuda, Takatomo Kano, Siddhant Arora, Marc Delcroix 외 arxiv

We investigate turn-taking in multimodal multi-party conversations using large language models (LLMs). We construct an evaluation framework for three tasks: addressee detection, turn-change prediction, and next speaker p…

'What are you referring to?' Evaluating the Ability of Multi-Modal Dialogue Models to Process Clarificational Exchanges

2023-07-28 · Javier Chiyah-Garcia, Alessandro Suglia, Arash Eshghi, Helen Hastie

Referential ambiguities arise in dialogue when a referring expression does not uniquely identify the intended referent for the addressee. Addressees usually detect such ambiguities immediately and work with the speaker t…

Referring Expression

A Multi-Modal Explainability Approach for Human-Aware Robots in Multi-Party Conversation

2024-05-20 · Iveta Bečková, Štefan Pócoš, Giulia Belgiovine, Marco Matarese 외

The addressee estimation (understanding to whom somebody is talking) is a fundamental task for human activity recognition in multi-party conversation scenarios. Specifically, in the field of human-robot interaction, it b…

Activity RecognitionBinary ClassificationExplainable artificial intelligenceHuman Activity Recognition

Do LLMs suffer from Multi-Party Hangover? A Diagnostic Approach to Addressee Recognition and Response Selection in Conversations

2024-09-27 · Nicolò Penzo, Maryam Sajedinia, Bruno Lepri, Sara Tonelli 외

Assessing the performance of systems to classify Multi-Party Conversations (MPC) is challenging due to the interconnection between linguistic and structural characteristics of conversations. Conventional evaluation metho…

Diagnostic