paper-with-me

홈 › Papers

Generalized Pose Space Embeddings for Training In-the-Wild using Anaylis-by-Synthesis

2024-11-13 · Dominik Borer, Jakob Buhmann, Martin Guay

Modern pose estimation models are trained on large, manually-labelled datasets which are costly and may not cover the full extent of human poses and appearances in the real world. With advances in neural rendering, analysis-by-synthesis and the ability to not only predict, but also render the pose, is becoming an appealing framework, which could alleviate the need for large scale manual labelling efforts. While recent work have shown the feasibility of this approach, the predictions admit many flips due to a simplistic intermediate skeleton representation, resulting in low precision and inhibiting the acquisition of any downstream knowledge such as three-dimensional positioning. We solve this problem with a more expressive intermediate skeleton representation capable of capturing the semantics of the pose (left and right), which significantly reduces flips. To successfully train this new representation, we extend the analysis-by-synthesis framework with a training protocol based on synthetic data. We show that our representation results in less flips and more accurate predictions. Our approach outperforms previous models trained with analysis-by-synthesis on standard benchmarks.

📄 PDF Abstract BibTeX arXiv:2411.08603

Code (0)

등록된 구현이 없습니다.

Tasks

Neural RenderingPose Estimation

Similar Papers 제목 키워드 기반

WildNet: Learning Domain Generalized Semantic Segmentation from the Wild

2022-04-04 · CVPR 2022 1 · Suhyeon Lee, Hongje Seong, Seongwon Lee, Euntai Kim

We present a new domain generalized semantic segmentation network named WildNet, which learns domain-generalized features by leveraging a variety of contents and styles from the wild. In domain generalization, the low ge…

Domain GeneralizationSemantic Segmentation

An Empirical Study and Analysis of Generalized Zero-Shot Learning for Object Recognition in the Wild

2016-05-13 · Wei-Lun Chao, Soravit Changpinyo, Boqing Gong, Fei Sha

Zero-shot learning (ZSL) methods have been studied in the unrealistic setting where test data are assumed to come from unseen classes only. In this paper, we advocate studying the problem of generalized zero-shot learnin…

Few-Shot LearningGeneralized Zero-Shot LearningObject RecognitionZero-Shot Learning

Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval

2026-01-30 · Ilyass Moummad, Marius Miron, David Robinson, Kawtar Zaher 외 arxiv

Large-scale biodiversity monitoring platforms increasingly rely on multimodal wildlife observations. While recent foundation models enable rich semantic representations across vision, audio, and language, retrieving rele…

parameter-efficient fine-tuningZero-shot GeneralizationImage Retrieval

AVGZSLNet: Audio-Visual Generalized Zero-Shot Learning by Reconstructing Label Features from Multi-Modal Embeddings

2020-05-27 · Pratik Mazumder, Pravendra Singh, Kranti Kumar Parida, Vinay P. Namboodiri

In this paper, we propose a novel approach for generalized zero-shot learning in a multi-modal setting, where we have novel classes of audio/video during testing that are not seen during training. We use the semantic rel…

DecoderGeneralized Zero-Shot LearningGZSL Video ClassificationRetrieval+3

On deep speaker embeddings for text-independent speaker recognition

2018-04-26 · Sergey Novoselov, Andrey Shulipa, Ivan Kremnev, Alexandr Kozlov 외

We investigate deep neural network performance in the textindependent speaker recognition task. We demonstrate that using angular softmax activation at the last classification layer of a classification neural network ins…

General ClassificationMetric LearningSpeaker RecognitionSpeaker Verification+1