paper-with-me

홈 › Papers

Learning When to Listen: Gated Affect Fusion for Human Motion Prediction

2026-07-01 · Jingni Huang arxiv

Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence of noisy multimodal observations. While facial affect potentially provides complementary behavioral cues, its practical utility and mechanistic boundaries within motion forecasting frameworks remain poorly understood. In this work, we present a systematic study investigating the utility and temporal limitations of affect-conditioned forecasting in-the-wild. We establish a rigorous multimodal pipeline combining MediaPipe body pose trajectories with HSEmotion facial affect representations, and introduce the Gated Affect Transformer (GAT) to dynamically regulate cross-modal information flow. Through extensive multi-horizon evaluations under a strict subject-wise protocol, we demonstrate that naive early cross-modal concatenation consistently degrades forecasting accuracy relative to pose-only baselines. Conversely, our proposed gating mechanism stabilizes cross-modal integration by adaptively controlling the affective stream. Crucially, controlled counterfactual experiments using shuffled and randomized affect inputs reveal that the learned gate successfully suppresses unstructured cross-modal noise while remaining responsive to plausible affective signals. Furthermore, our empirical results indicate that facial affect features provide bounded, horizon-dependent predictive cues strictly within short-to-medium windows (e.g., 30 frames), whereas long-term trajectories remain predominantly governed by intrinsic kinematic continuity. Our findings provide empirical evidence that facial affect should be regarded as a complementary behavioral cue rather than a dominant driver of future motion, offering practical guidance for selective multimodal fusion in unconstrained human motion forecasting.

📄 PDF Abstract BibTeX arXiv:2607.00296

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Forecasting

Similar Papers 제목 키워드 기반

The role of cue enhancement and frequency fine-tuning in hearing impaired phone recognition

2019-08-09 · Ali Abavisani, Mark A. Hasegawa-Johnson

A speech-based hearing test is designed to identify the susceptible error-prone phones for individual hearing impaired (HI) ear. Only robust tokens in the experiment noise levels had been chosen for the test. The noise-r…

Speaker discrimination in humans and machines: Effects of speaking style variability

2020-08-08 · Amber Afshan, Jody Kreiman, Abeer Alwan

Does speaking style variation affect humans' ability to distinguish individuals from their voices? How do humans compare with automatic systems designed to discriminate between voices? In this paper, we attempt to answer…

Speaker Verification

Utilizing Mood-Inducing Background Music in Human-Robot Interaction

2023-08-28 · Elad Liebman, Peter Stone

Past research has clearly established that music can affect mood and that mood affects emotional and cognitive processing, and thus decision-making. It follows that if a robot interacting with a person needs to predict t…

Decision Making

Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots

2026-08-17 · Zi Haur Pang, Casey Kennington, Tatsuya Kawahara arxiv

Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primaril…

Response Generation

The OMG-Empathy Dataset: Evaluating the Impact of Affective Behavior in Storytelling

2019-08-30 · Pablo Barros, Nikhil Churamani, Angelica Lim, Stefan Wermter

Processing human affective behavior is important for developing intelligent agents that interact with humans in complex interaction scenarios. A large number of current approaches that address this problem focus on class…