paper-with-me

홈 › Papers

Remembering What Is Important: A Factorised Multi-Head Retrieval and Auxiliary Memory Stabilisation Scheme for Human Motion Prediction

2023-05-19 · Tharindu Fernando, Harshala Gammulle, Sridha Sridharan, Simon Denman, Clinton Fookes

Humans exhibit complex motions that vary depending on the task that they are performing, the interactions they engage in, as well as subject-specific preferences. Therefore, forecasting future poses based on the history of the previous motions is a challenging task. This paper presents an innovative auxiliary-memory-powered deep neural network framework for the improved modelling of historical knowledge. Specifically, we disentangle subject-specific, task-specific, and other auxiliary information from the observed pose sequences and utilise these factorised features to query the memory. A novel Multi-Head knowledge retrieval scheme leverages these factorised feature embeddings to perform multiple querying operations over the historical observations captured within the auxiliary memory. Moreover, our proposed dynamic masking strategy makes this feature disentanglement process dynamic. Two novel loss functions are introduced to encourage diversity within the auxiliary memory while ensuring the stability of the memory contents, such that it can locate and store salient information that can aid the long-term prediction of future motion, irrespective of data imbalances or the diversity of the input data distribution. With extensive experiments conducted on two public benchmarks, Human3.6M and CMU-Mocap, we demonstrate that these design choices collectively allow the proposed approach to outperform the current state-of-the-art methods by significant margins: $>$ 17\% on the Human3.6M dataset and $>$ 9\% on the CMU-Mocap dataset.

📄 PDF Abstract BibTeX arXiv:2305.11394

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementDiversityHuman motion predictionHuman Pose Forecastingmotion predictionRetrieval

Similar Papers 제목 키워드 기반

From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

2026-06-26 · Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Quang Minh Nguyen 외 arxiv

Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by learning uncertainty from internal model signals. Yet, recent methods vary simultane…

Non-parametric Memory for Spatio-Temporal Segmentation of Construction Zones for Self-Driving

2021-01-18 · Min Bai, Shenlong Wang, Kelvin Wong, Ersin Yumer 외

In this paper, we introduce a non-parametric memory representation for spatio-temporal segmentation that captures the local space and time around an autonomous vehicle (AV). Our representation has three important propert…

Modernising Reinforcement Learning-Based Navigation for Embodied Semantic Scene Graph Generation

2026-03-26 · Roman Küble, Marco Hüller, Mrunmai Phatak, Rainer Lienhart 외 arxiv

Semantic world models enable embodied agents to reason about objects, relations, and spatial context beyond purely geometric representations. In Organic Computing, such models are a key enabler for objective-driven self-…

Reinforcement LearningScene Graph Generation

Cross Temporal Recurrent Networks for Ranking Question Answer Pairs

2017-11-21 · Yi Tay, Luu Anh Tuan, Siu Cheung Hui

Temporal gates play a significant role in modern recurrent-based neural encoders, enabling fine-grained control over recursive compositional operations over time. In recurrent models such as the long short-term memory (L…

How Does Attention Work in Vision Transformers? A Visual Analytics Attempt

2023-03-24 · Yiran Li, Junpeng Wang, Xin Dai, Liang Wang 외

Vision transformer (ViT) expands the success of transformer models from sequential data to images. The model decomposes an image into many smaller patches and arranges them into a sequence. Multi-head self-attentions are…