paper-with-me

홈 › Papers

Image-free Domain Generalization via CLIP for 3D Hand Pose Estimation

2022-10-30 · Seongyeong Lee, Hansoo Park, Dong Uk Kim, Jihyeon Kim, Muhammadjon Boboev, Seungryul Baek

RGB-based 3D hand pose estimation has been successful for decades thanks to large-scale databases and deep learning. However, the hand pose estimation network does not operate well for hand pose images whose characteristics are far different from the training data. This is caused by various factors such as illuminations, camera angles, diverse backgrounds in the input images, etc. Many existing methods tried to solve it by supplying additional large-scale unconstrained/target domain images to augment data space; however collecting such large-scale images takes a lot of labors. In this paper, we present a simple image-free domain generalization approach for the hand pose estimation framework that uses only source domain data. We try to manipulate the image features of the hand pose estimation network by adding the features from text descriptions using the CLIP (Contrastive Language-Image Pre-training) model. The manipulated image features are then exploited to train the hand pose estimation network via the contrastive learning framework. In experiments with STB and RHD datasets, our algorithm shows improved performance over the state-of-the-art domain generalization approaches.

📄 PDF Abstract BibTeX arXiv:2210.16788

Code (0)

등록된 구현이 없습니다.

Tasks

3D Hand Pose EstimationContrastive LearningDomain GeneralizationHand Pose EstimationPose Estimation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

DomainVerse: A Benchmark Towards Real-World Distribution Shifts For Tuning-Free Adaptive Domain Generalization

2024-03-05 · Feng Hou, Jin Yuan, Ying Yang, Yang Liu 외

Traditional cross-domain tasks, including domain adaptation and domain generalization, rely heavily on training model by source domain data. With the recent advance of vision-language models (VLMs), viewed as natural sou…

Domain AdaptationDomain Generalization

AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection

2025-05-15 · Bin-Bin Gao, Yue Zhou, Jiangtao Yan, Yuezhi Cai 외

Universal visual anomaly detection aims to identify anomalies from novel or unseen vision domains without additional fine-tuning, which is critical in open scenarios. Recent studies have demonstrated that pre-trained vis…

Anomaly Detection

Exploring the Transfer Learning Capabilities of CLIP in Domain Generalization for Diabetic Retinopathy

2023-08-27 · Sanoojan Baliah, Fadillah A. Maani, Santosh Sanjeev, Muhammad Haris Khan

Diabetic Retinopathy (DR), a leading cause of vision impairment, requires early detection and treatment. Developing robust AI models for DR classification holds substantial potential, but a key challenge is ensuring thei…

ClassificationDomain GeneralizationTransfer Learning

In the Era of Prompt Learning with Vision-Language Models

2024-11-07 · Ankit Jha

Large-scale foundation models like CLIP have shown strong zero-shot generalization but struggle with domain shifts, limiting their adaptability. In our work, we introduce \textsc{StyLIP}, a novel domain-agnostic prompt l…

Domain AdaptationDomain GeneralizationPrompt LearningSemantic Segmentation+2

When Semantics Regulate: Rethinking Patch Shuffle and Internal Bias for Generated Image Detection with CLIP

2025-11-24 · Beilin Chu, Weike You, Mengtao Li, Tingting Zheng 외 arxiv

The rapid progress of GANs and Diffusion Models poses new challenges for detecting AI-generated images. Although CLIP-based detectors exhibit promising generalization, they often rely on semantic cues rather than generat…

Domain Generalization