paper-with-me

홈 › Papers

L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream

2024-01-01 · CVPR 2024 1 · Jingtao Sun, Yaonan Wang, Mingtao Feng, Yulan Guo, Ajmal Mian, Mike Zheng Shou

3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper we aim to handle a real-time multi-task for 6-DoF pose tracking of unknown objects leveraging 3D-language pre-training scheme from a series of 3D point cloud video streams while simultaneously performing 3D shape reconstruction in current observation. To this end we present a generic Language-to-4D modeling paradigm termed L4D-Track that tackles zero-shot 6-DoF \underline Track ing and shape reconstruction by learning pairwise implicit 3D representation and multi-level multi-modal alignment. Our method constitutes two core parts. 1) Pairwise Implicit 3D Space Representation that establishes spatial-temporal to language coherence descriptions across continuous 3D point cloud video. 2) Language-to-4D Association and Contrastive Alignment enables multi-modality semantic connections between 3D point cloud video and language. Our method trained exclusively on public NOCS-REAL275 dataset achieves promising results on both two publicly benchmarks. This not only shows powerful generalization performance but also proves its remarkable capability in zero-shot inference.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D Shape ReconstructionPose Tracking

Similar Papers 제목 키워드 기반

Online Adaptation for Implicit Object Tracking and Shape Reconstruction in the Wild

2021-11-24 · Jianglong Ye, Yuntao Chen, Naiyan Wang, Xiaolong Wang

Tracking and reconstructing 3D objects from cluttered scenes are the key components for computer vision, robotics and autonomous driving systems. While recent progress in implicit function has shown encouraging results o…

3D Shape ReconstructionAutonomous DrivingObject Tracking

Robust Performance-driven 3D Face Tracking in Long Range Depth Scenes

2015-07-10 · Hai X. Pham, Chongyu Chen, Luc N. Dao, Vladimir Pavlovic 외

We introduce a novel robust hybrid 3D face tracking framework from RGBD video streams, which is capable of tracking head pose and facial actions without pre-calibration or intervention from a user. In particular, we emph…

3D ReconstructionFace Model

DoubleFusion: Real-time Capture of Human Performances with Inner Body Shapes from a Single Depth Sensor

2018-04-17 · CVPR 2018 6 · Tao Yu, Zerong Zheng, Kaiwen Guo, Jianhui Zhao 외

We propose DoubleFusion, a new real-time system that combines volumetric dynamic reconstruction with data-driven template fitting to simultaneously reconstruct detailed geometry, non-rigid motion and the inner human body…

Dynamic Reconstruction

InterTrack: Tracking Human Object Interaction without Object Templates

2024-08-25 · Xianghui Xie, Jan Eric Lenssen, Gerard Pons-Moll

Tracking human object interaction from videos is important to understand human behavior from the rapidly growing stream of video data. Previous video-based methods require predefined object templates while single-image-b…

Human-Object Interaction DetectionObjectPose Tracking

LatentHuman: Shape-and-Pose Disentangled Latent Representation for Human Bodies

2021-11-30 · Sandro Lombardi, Bangbang Yang, Tianxing Fan, Hujun Bao 외

3D representation and reconstruction of human bodies have been studied for a long time in computer vision. Traditional methods rely mostly on parametric statistical linear models, limiting the space of possible bodies to…

3D Reconstructionmotion retargetingPose Tracking