paper-with-me

홈 › Papers

Look, Listen and Learn - A Multimodal LSTM for Speaker Identification

2016-02-13 · Jimmy Ren, Yongtao Hu, Yu-Wing Tai, Chuan Wang, Li Xu, Wenxiu Sun, Qiong Yan

Speaker identification refers to the task of localizing the face of a person who has the same identity as the ongoing voice in a video. This task not only requires collective perception over both visual and auditory signals, the robustness to handle severe quality degradations and unconstrained content variations are also indispensable. In this paper, we describe a novel multimodal Long Short-Term Memory (LSTM) architecture which seamlessly unifies both visual and auditory modalities from the beginning of each sequence input. The key idea is to extend the conventional LSTM by not only sharing weights across time steps, but also sharing weights across modalities. We show that modeling the temporal dependency across face and voice can significantly improve the robustness to content quality degradations and variations. We also found that our multimodal LSTM is robustness to distractors, namely the non-speaking identities. We applied our multimodal LSTM to The Big Bang Theory dataset and showed that our system outperforms the state-of-the-art systems in speaker identification with lower false alarm rate and higher recognition accuracy.

📄 PDF Abstract BibTeX arXiv:1602.04364

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Identification

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

A Multimodal LSTM for Predicting Listener Empathic Responses Over Time

2018-12-12 · Zhi-Xuan Tan, Arushi Goel, Thanh-Son Nguyen, Desmond C. Ong

People naturally understand the emotions of-and often also empathize with-those around them. In this paper, we predict the emotional valence of an empathic listener over time as they listen to a speaker narrating a life …

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

2026-08-17 · Zi Haur Pang, Casey Kennington, Tatsuya Kawahara arxiv

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primaril…

Response Generation

MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion Model

2023-08-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlook…

DenoisingDiversity

Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion

2022-04-18 · CVPR 2022 1 · Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li 외

We present a framework for modeling interactional communication in dyadic conversations: given multimodal inputs of a speaker, we autoregressively output multiple possibilities of corresponding listener motion. We combin…

Read, Look or Listen? What's Needed for Solving a Multimodal Dataset

2023-07-06 · Netta Madvil, Yonatan Bitton, Roy Schwartz

The prevalence of large-scale multimodal datasets presents unique challenges in assessing dataset quality. We propose a two-step method to analyze multimodal datasets, which leverages a small seed of human annotation to …

Question AnsweringSpeaker IdentificationVideo Question Answering