paper-with-me

홈 › Papers

Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from Videos

2023-08-20 · ICCV 2023 1 · Haoyuan Li, Haoye Dong, Hanchao Jia, Dong Huang, Michael C. Kampffmeyer, Liang Lin, Xiaodan Liang

Multi-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person detection and tracking stages are performed in a multi-person setting, while temporal dynamics are only modeled for one person at a time. Consequently, their performance is severely limited by the lack of inter-person interactions in the spatial-temporal mesh recovery, as well as by detection and tracking defects. To address these challenges, we propose the Coordinate transFormer (CoordFormer) that directly models multi-person spatial-temporal relations and simultaneously performs multi-mesh recovery in an end-to-end manner. Instead of partitioning the feature map into coarse-scale patch-wise tokens, CoordFormer leverages a novel Coordinate-Aware Attention to preserve pixel-level spatial-temporal coordinate information. Additionally, we propose a simple, yet effective Body Center Attention mechanism to fuse position information. Extensive experiments on the 3DPW dataset demonstrate that CoordFormer significantly improves the state-of-the-art, outperforming the previously best results by 4.2%, 8.8% and 4.7% according to the MPJPE, PAMPJPE, and PVE metrics, respectively, while being 40% faster than recent video-based approaches. The released code can be found at https://github.com/Li-Hao-yuan/CoordFormer.

📄 PDF Abstract BibTeX arXiv:2308.10334

Code (0)

등록된 구현이 없습니다.

Tasks

Human Detection

Similar Papers 제목 키워드 기반

RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation

2023-12-12 · CVPR 2024 1 · Peng Lu, Tao Jiang, Yining Li, Xiangtai Li 외

Real-time multi-person pose estimation presents significant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases, existing one-stage metho…

GPUMulti-Person Pose EstimationPose Estimation

YORO -- Lightweight End to End Visual Grounding

2022-11-15 · Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani, R. Manmatha 외

We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural language. Unlike the recent trend in th…

Natural Language QueriesVisual Grounding

End-to-End Spatial-Temporal Transformer for Real-time 4D HOI Reconstruction

2026-03-15 · Haoyu Zhang, Wei Zhai, Yuhang Yang, Yang Cao 외 arxiv

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing metho…

CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition

2025-08-31 · Yusen Peng, Alper Yilmaz arxiv

Skeleton-based human action recognition leverages sequences of human joint coordinates to identify actions performed in videos. Owing to the intrinsic spatiotemporal structure of skeleton data, Graph Convolutional Networ…

Representation LearningAction ClassificationAction Recognition

Coordinate In and Value Out: Training Flow Transformers in Ambient Space

2024-12-05 · Yuyang Wang, Anurag Ranjan, Josh Susskind, Miguel Angel Bautista

Flow matching models have emerged as a powerful method for generative modeling on domains like images or videos, and even on unstructured data like 3D point clouds. These models are commonly trained in two stages: first,…