paper-with-me

홈 › Papers

Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular Video

2022-03-16 · CVPR 2022 1 · Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao

Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature representations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/.

📄 PDF Abstract BibTeX arXiv:2203.08534

Code (0)

등록된 구현이 없습니다.

Tasks

3D human pose and shape estimation3D Human Pose Estimation

Similar Papers 제목 키워드 기반

Investigating Pose Representations and Motion Contexts Modeling for 3D Motion Prediction

2021-12-30 · Zhenguang Liu, Shuang Wu, Shuyuan Jin, Shouling Ji 외

Predicting human motion from historical pose sequence is crucial for a machine to succeed in intelligent interactions with humans. One aspect that has been obviated so far, is the fact that how we represent the skeletal …

motion predictionPrediction

To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions

2024-03-19 · Daniel Tanneberg, Felix Ocker, Stephan Hasler, Joerg Deigmoeller 외

How can a robot provide unobtrusive physical support within a group of humans? We present Attentive Support, a novel interaction concept for robots to support a group of humans. It combines scene perception, dialogue acq…

Common Sense Reasoning

DreamPose3D: Hallucinative Diffusion with Prompt Learning for 3D Human Pose Estimation

2025-11-12 · Jerrin Bright, Yuhao Chen, John S. Zelek arxiv

Accurate 3D human pose estimation remains a critical yet unresolved challenge, requiring both temporal coherence across frames and fine-grained modeling of joint relationships. However, most existing methods rely solely …

3D Human Pose Estimation3D Pose Estimation

GRIP: Generating Interaction Poses Using Spatial Cues and Latent Consistency

2023-08-22 · Omid Taheri, Yi Zhou, Dimitrios Tzionas, Yang Zhou 외

Hands are dexterous and highly versatile manipulators that are central to how humans interact with objects and their environment. Consequently, modeling realistic hand-object interactions, including the subtle motion of …

Mixed RealityObject

A Self-Attentive Emotion Recognition Network

2019-04-24 · Harris Partaourides, Kostantinos Papadamou, Nicolas Kourtellis, Ilias Leontiadis 외

Modern deep learning approaches have achieved groundbreaking performance in modeling and classifying sequential data. Specifically, attention networks constitute the state-of-the-art paradigm for capturing long temporal …

DecoderEmotion Recognition