paper-with-me

Papers

Non-Local Latent Relation Distillation for Self-Adaptive 3D Human Pose Estimation

2022-04-05 · NeurIPS 2021 12 · Jogendra Nath Kundu, Siddharth Seth, Anirudh Jamkhandi, Pradyumna YM, Varun Jampani, Anirban Chakraborty, R. Venkatesh Babu

Available 3D human pose estimation approaches leverage different forms of strong (2D/3D pose) or weak (multi-view or depth) paired supervision. Barring synthetic or in-studio domains, acquiring such supervision for each new target environment is highly inconvenient. To this end, we cast 3D pose learning as a self-supervised adaptation problem that aims to transfer the task knowledge from a labeled source domain to a completely unpaired target. We propose to infer image-to-pose via two explicit mappings viz. image-to-latent and latent-to-pose where the latter is a pre-learned decoder obtained from a prior-enforcing generative adversarial auto-encoder. Next, we introduce relation distillation as a means to align the unpaired cross-modal samples i.e. the unpaired target videos and unpaired 3D pose sequences. To this end, we propose a new set of non-local relations in order to characterize long-range latent pose interactions unlike general contrastive relations where positive couplings are limited to a local neighborhood structure. Further, we provide an objective way to quantify non-localness in order to select the most effective relation set. We evaluate different self-adaptation settings and demonstrate state-of-the-art 3D human pose estimation performance on standard benchmarks.

📄 PDF Abstract BibTeX arXiv:2204.01971

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human Pose EstimationDecoderPose EstimationRelationUnsupervised 3D Human Pose EstimationWeakly-supervised 3D Human Pose Estimation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

2026-07-28 · Yuqi Li, Yi-Cheng Lin, Xianglong Wang, Kuo Yang 외 arxiv

On-device speech emotion recognition (SER) is critical for real-time applications, yet large self-supervised models that excel at SER are too costly for edge devices. Multi-teacher knowledge distillation can compress the…

Speech Emotion RecognitionKnowledge Distillation

Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning

2023-03-23 · ICCV 2023 1 · Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran 외

Most self-supervised methods for representation learning leverage a cross-view consistency objective i.e., they maximize the representation similarity of a given image's augmented views. Recent work NNCLR goes beyond the…

Contrastive LearningRepresentation Learning

AURORA: Contextual Orthogonalization for Geometric Representation Learning in Healthcare Foundation Models

2026-05-18 · Yuanyun Zhang, Shi Li arxiv

Recent healthcare foundation models have achieved strong predictive performance through large scale self supervised learning, yet their latent representations frequently entangle physiologic severity, intervention intens…

Representation Learning

Projected Latent Distillation for Data-Agnostic Consolidation in Distributed Continual Learning

2023-03-28 · Antonio Carta, Andrea Cossu, Vincenzo Lomonaco, Davide Bacciu 외

Distributed learning on the edge often comprises self-centered devices (SCD) which learn local tasks independently and are unwilling to contribute to the performance of other SDCs. How do we achieve forward transfer at z…

Continual LearningKnowledge Distillation

SCALAR: Self-Calibrating Adaptive Latent Attention Representation Learning

2025-10-18 · Farwa Abbas, Hussain Ahmad, Claudia Szabo arxiv

High-dimensional, heterogeneous data with complex feature interactions pose significant challenges for traditional predictive modeling approaches. While Projection to Latent Structures (PLS) remains a popular technique, …

Representation Learning