paper-with-me

홈 › Papers

V-VIPE: Variational View Invariant Pose Embedding

2024-07-09 · Mara Levy, Abhinav Shrivastava

Learning to represent three dimensional (3D) human pose given a two dimensional (2D) image of a person, is a challenging problem. In order to make the problem less ambiguous it has become common practice to estimate 3D pose in the camera coordinate space. However, this makes the task of comparing two 3D poses difficult. In this paper, we address this challenge by separating the problem of estimating 3D pose from 2D images into two steps. We use a variational autoencoder (VAE) to find an embedding that represents 3D poses in canonical coordinate space. We refer to this embedding as variational view-invariant pose embedding V-VIPE. Using V-VIPE we can encode 2D and 3D poses and use the embedding for downstream tasks, like retrieval and classification. We can estimate 3D poses from these embeddings using the decoder as well as generate unseen 3D poses. The variability of our encoding allows it to generalize well to unseen camera views when mapping from 2D space. To the best of our knowledge, V-VIPE is the only representation to offer this diversity of applications. Code and more information can be found at https://v-vipe.github.io/.

📄 PDF Abstract BibTeX arXiv:2407.07092

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDiversity

Similar Papers 제목 키워드 기반

KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM

2025-12-01 · Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov 외 arxiv

We present KM-ViPE (Knowledge Mapping Video Pose Engine), a real-time open-vocabulary SLAM framework for uncalibrated monocular cameras in dynamic environments. Unlike systems requiring depth sensors and offline calibrat…

Semantic SLAM

Learning Mid-level Filters for Person Re-identification

2014-06-01 · CVPR 2014 6 · Rui Zhao, Wanli Ouyang, Xiaogang Wang

In this paper, we propose a novel approach of learning mid-level filters from automatically discovered patch clusters for person re-identification. It is well motivated by our study on what are good filters for person re…

ClusteringPatch MatchingPerson Re-Identification

Learning Invariant Color Features for Person Re-Identification

2014-10-04 · Rahul Rama Varior, Gang Wang, Jiwen Lu

Matching people across multiple camera views known as person re-identification, is a challenging problem due to the change in visual appearance caused by varying lighting conditions. The perceived color of the subject ap…

Person Re-Identification

Pose Invariant Embedding for Deep Person Re-identification

2017-01-26 · Liang Zheng, Yujia Huang, Huchuan Lu, Yi Yang

Pedestrian misalignment, which mainly arises from detector errors and pose variations, is a critical problem for a robust person re-identification (re-ID) system. With bad alignment, the background noise will significant…

Person Re-IdentificationPose EstimationRetrieval

RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments

2026-04-28 · Zaid Nasser, Mikhail Iumanov, Tianhao Li, Maxim Popov 외 arxiv

We present RADIO-ViPE (Reduce All Domains Into One -- Video Pose Engine), an online semantic SLAM system that enables geometry-aware open-vocabulary grounding, associating arbitrary natural language queries with localize…

Natural Language QueriesSemantic SLAM