paper-with-me

홈 › Papers

SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation

2025-01-16 · Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Chen Wei, Qingping Sun, Haiyi Mei, Yanjun Wang, Hui En Pang, Mingyuan Zhang, Lei Zhang, Chen Change Loy, Atsushi Yamashita, Lei Yang, Ziwei Liu

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods focus on training innovative architectural designs on confined datasets. In this work, we investigate the impact of scaling up EHPS towards a family of generalist foundation models. 1) For data scaling, we perform a systematic investigation on 40 EHPS datasets, encompassing a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. Ultimately, we achieve diminishing returns at 10M training instances from diverse data sources. 2) For model scaling, we take advantage of vision transformers (up to ViT-Huge as the backbone) to study the scaling law of model sizes in EHPS. To exclude the influence of algorithmic design, we base our experiments on two minimalist architectures: SMPLer-X, which consists of an intermediate step for hand and face localization, and SMPLest-X, an even simpler version that reduces the network to its bare essentials and highlights significant advances in the capture of articulated hands. With big data and the large model, the foundation models exhibit strong performance across diverse test benchmarks and excellent transferability to even unseen environments. Moreover, our finetuning strategy turns the generalist into specialist models, allowing them to achieve further performance boosts. Notably, our foundation models consistently deliver state-of-the-art results on seven benchmarks such as AGORA, UBody, EgoBody, and our proposed SynHand dataset for comprehensive hand evaluation. (Code is available at: https://github.com/wqyin/SMPLest-X).

📄 PDF Abstract BibTeX arXiv:2501.09782

Code (3)

openxrlab/xrfeitoria 공식 구현
wqyin/smplest-x 공식 구현 pytorch
caizhongang/SMPLer-X pytorch

Tasks

Benchmarking

Methods 이 논문이 사용한 방법론

BASE 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

SMART: SMPLest-X Mesh Adaptation and RAFT Tracking for Soccer Pose Estimation

2026-05-29 · Parthsarthi Rawat arxiv

We present our approach to the FIFA Skeletal Tracking Challenge 2026, which requires estimating 3D world-space poses of soccer players from broadcast video. Our method finetunes SMPLest-X (ViT-H, 687 M parameters) via a …

Pose Estimation

Expressive Power of Implicit Models: Rich Equilibria and Test-Time Scaling

2025-10-04 · Jialin Liu, Lisang Ding, Stanley Osher, Wotao Yin arxiv

Implicit models, an emerging model class, compute outputs by iterating a single parameter block to a fixed point. This architecture realizes an infinite-depth, weight-tied network that trains with constant memory, signif…

Image Reconstruction

Scaling Behavior Foundation Model for Humanoid Robots

2026-07-16 · Weishuai Zeng, Kangning Yin, Xiaojie Niu, Shunlin Lu 외 arxiv

Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across diverse environmental contexts, making it a cornerstone for generalist embodied …

SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation

2023-09-29 · NeurIPS 2023 11 · Zhongang Cai, Wanqi Yin, Ailing Zeng, Chen Wei 외

Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined…

3D Human Pose Estimation3D Human Reconstruction3D Multi-Person Mesh RecoveryBenchmarking

Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller

2024-03-12 · Chuanqi Zang, Jiji Tang, Rongsheng Zhang, Zeng Zhao 외

Storytelling aims to generate reasonable and vivid narratives based on an ordered image stream. The fidelity to the image story theme and the divergence of story plots attract readers to keep reading. Previous works iter…

Story Generation