paper-with-me

Papers

CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks

2026-01-19 · Mingshuang Luo, Ruibing Hou, Bo Chao, Hong Chang, Zimo Liu, Yaowei Wang, Shiguang Shan arxiv

Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale unlabeled human image datasets, there is an increasing need for a general unsupervised pre-training model capable of supporting diverse human-centric downstream tasks. To achieve this goal, we propose CLASP (CLIP-guided Adaptable Self-suPervised learning), a novel framework designed for unsupervised pre-training in human-centric visual tasks. CLASP leverages the powerful vision-language model CLIP to generate both low-level (e.g., body parts) and high-level (e.g., attributes) semantic pseudo-labels. These multi-level semantic cues are then integrated into the learned visual representations, enriching their expressiveness and generalizability. Recognizing that different downstream tasks demand varying levels of semantic granularity, CLASP incorporates a Prompt-Controlled Mixture-of-Experts (MoE) module. MoE dynamically adapts feature extraction based on task-specific prompts, mitigating potential feature conflicts and enhancing transferability. Furthermore, CLASP employs a multi-task pre-training strategy, where part- and attribute-level pseudo-labels derived from CLIP guide the representation learning process. Extensive experiments across multiple benchmarks demonstrate that CLASP consistently outperforms existing unsupervised pre-training methods, advancing the field of human-centric visual analysis.

📄 PDF Abstract BibTeX arXiv:2601.13133

Code (0)

등록된 구현이 없습니다.

Tasks

Unsupervised Pre-trainingSelf-Supervised LearningRepresentation Learning

Similar Papers 제목 키워드 기반

CLIP-S$^4$: Language-Guided Self-Supervised Semantic Segmentation

2023-05-01 · Wenbin He, Suphanut Jamonnak, Liang Gou, Liu Ren

Existing semantic segmentation approaches are often limited by costly pixel-wise annotations and predefined classes. In this work, we present CLIP-S$^4$ that leverages self-supervised pixel representation learning and vi…

Contrastive LearningLanguage ModellingRepresentation LearningSegmentation+2

CLIP-S4: Language-Guided Self-Supervised Semantic Segmentation

2023-01-01 · CVPR 2023 1 · Wenbin He, Suphanut Jamonnak, Liang Gou, Liu Ren

Existing semantic segmentation approaches are often limited by costly pixel-wise annotations and predefined classes. In this work, we present CLIP-S^4 that leverages self-supervised pixel representation learning and …

Contrastive LearningLanguage ModellingRepresentation LearningSegmentation+2

LOC: A General Language-Guided Framework for Open-Set 3D Occupancy Prediction

2025-10-25 · Yuhang Gao, Xiang Xiang, Sheng Zhong, Guoyou Wang arxiv

Vision-Language Models (VLMs) have shown significant progress in open-set challenges. However, the limited availability of 3D datasets hinders their effective application in 3D scene understanding. We propose LOC, a gene…

Self-Supervised LearningContrastive LearningSemantic SimilarityScene Understanding

Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning

2023-03-02 · Bo Wan, Yongfei Liu, Desen Zhou, Tinne Tuytelaars 외

Human object interaction (HOI) detection plays a crucial role in human-centric scene understanding and serves as a fundamental building-block for many vision tasks. One generalizable and scalable strategy for HOI detecti…

Human-Object Interaction DetectionKnowledge DistillationObjectRepresentation Learning+1

ClipFace: Text-guided Editing of Textured 3D Morphable Models

2022-12-02 · Shivangi Aneja, Justus Thies, Angela Dai, Matthias Nießner

We propose ClipFace, a novel self-supervised approach for text-guided editing of textured 3D morphable model of faces. Specifically, we employ user-friendly language prompts to enable control of the expressions as well a…

Texture Synthesis