paper-with-me

Papers

MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation

2025-01-01 · CVPR 2025 1 · Aviral Chharia, Wenbo Gou, Haoye Dong

While significant progress has been made in single-view 3D human pose estimation, multi-view 3D human pose estimation remains challenging, particularly in terms of generalizing to new camera configurations. Existing attention-based transformers often struggle to accurately model the spatial arrangement of keypoints, especially in occluded scenarios. Additionally, they tend to overfit specific camera arrangements and visual scenes from training data, resulting in substantial performance drops in new settings. In this study, we introduce a novel Multi-View State Space Modeling framework, named MV-SSM, for robustly estimating 3D human keypoints. We explicitly model the joint spatial sequence at two distinct levels: the feature level from multi-view images and the person keypoint level. We propose a Projective State Space (PSS) block to learn a generalized representation of joint spatial arrangements using state space modeling. Moreover, we modify Mamba's traditional scanning into an effective Grid Token-guided Bidirectional Scanning (GTBS), which is integral to the PSS block. Multiple experiments demonstrate that MV-SSM achieves strong generalization, outperforming state-of-the-art methods: +10.8 on AP25 on the challenging three-camera setting in CMU Panoptic, +7.0 on AP25 on varying camera arrangements, and +15.3 PCP on Campus A1 in cross-dataset evaluations. Project Website: https://aviralchharia.github.io/MV-SSM

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human Pose EstimationMulti-view 3D Human Pose EstimationPose Estimation

Similar Papers 제목 키워드 기반

MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation

2025-08-31 · Aviral Chharia, Wenbo Gou, Haoye Dong arxiv

While significant progress has been made in single-view 3D human pose estimation, multi-view 3D human pose estimation remains challenging, particularly in terms of generalizing to new camera configurations. Existing atte…

3D Human Pose Estimation

DCHM: Depth-Consistent Human Modeling for Multiview Detection

2025-07-19 · Jiahao Ma, Tianyu Wang, Miaomiao Liu, David Ahmedt-Aristizabal 외 arxiv

Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for …

Pedestrian DetectionMultiview DetectionDepth EstimationPoint Clouds

S5 Framework: A Review of Self-Supervised Shared Semantic Space Optimization for Multimodal Zero-Shot Learning

2022-01-16 · ACL ARR January 2022 1 · Anonymous

In this review, we aim to inspire research into Self-Supervised Shared Semantic Space (S5) multimodal learning problems. We equip non-expert researchers with a framework of informed modeling decisions via an extensive li…

DenoisingZero-Shot Learning

UV Gaussians: Joint Learning of Mesh Deformation and Gaussian Textures for Human Avatar Modeling

2024-03-18 · Yujiao Jiang, Qingmin Liao, Xiaoyu Li, Li Ma 외

Reconstructing photo-realistic drivable human avatars from multi-view image sequences has been a popular and challenging topic in the field of computer vision and graphics. While existing NeRF-based methods can achieve h…

NeRF

One-Shot Feed-Forward 360$^{\circ}$ Animatable Avatar via Inpainted UV-Space Gaussian Modeling

2026-01-19 · Shuling Zhao, Dan Xu arxiv

Building one-shot 3D animatable head avatars is an important yet challenging problem. Existing methods generally collapse under large camera pose variations, compromising the realism of 3D avatars. In this work, we propo…

3D Reconstruction