paper-with-me

홈 › Papers

MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition

2023-11-30 · Dan Song, Xinwei Fu, Ning Liu, Weizhi Nie, Wenhui Li, Lanjun Wang, You Yang, AnAn Liu

Large-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models for 3D shapes, recent methods utilize language-image pre-training to realize zero-shot 3D shape recognition. However, due to the modality gap, pretrained language-image models are not confident enough in the generalization to 3D shape recognition. Consequently, this paper aims to improve the confidence with view selection and hierarchical prompts. Leveraging the CLIP model as an example, we employ view selection on the vision side by identifying views with high prediction confidence from multiple rendered views of a 3D shape. On the textual side, the strategy of hierarchical prompts is proposed for the first time. The first layer prompts several classification candidates with traditional class-level descriptions, while the second layer refines the prediction based on function-level descriptions or further distinctions between the candidates. Remarkably, without the need for additional training, our proposed method achieves impressive zero-shot 3D classification accuracies of 84.44%, 91.51%, and 66.17% on ModelNet40, ModelNet10, and ShapeNet Core55, respectively. Furthermore, we will make the code publicly available to facilitate reproducibility and further research in this area.

📄 PDF Abstract BibTeX arXiv:2311.18402

Code (0)

등록된 구현이 없습니다.

Tasks

3D Classification3D Shape RecognitionZero-shot 3D classification

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

PointCLIP: Point Cloud Understanding by CLIP

2021-12-04 · CVPR 2022 1 · Renrui Zhang, Ziyu Guo, Wei zhang, Kunchang Li 외

Recently, zero-shot and few-shot learning via Contrastive Vision-Language Pre-training (CLIP) have shown inspirational performance on 2D visual recognition, which learns to match images with their corresponding texts in …

3D Open-Vocabulary Instance SegmentationFew-Shot LearningOpen Vocabulary Object DetectionTraining-free 3D Part Segmentation+5

CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation

2024-06-21 · Muhammad Ali, Salman Khan

Multi-label classification is an essential task utilized in a wide variety of real-world applications. Multi-label zero-shot learning is a method for classifying images into multiple unseen categories for which no traini…

ClassificationDecoderGeneralized Zero-Shot LearningMulti-Label Classification+5

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

2024-06-02 · Yunheng Li, Zhongyu Li, Quansheng Zeng, Qibin Hou 외

Pre-trained vision-language models, e.g., CLIP, have been successfully applied to zero-shot semantic segmentation. Existing CLIP-based approaches primarily utilize visual features from the last layer to align with text e…

SegmentationSemantic SegmentationZero-Shot Semantic Segmentation

CRoF: CLIP-based Robust Few-shot Learning on Noisy Labels

2024-12-17 · Shizhuo Deng, Bowen Han, Jiaqi Chen, Hao Wang 외

Noisy labels threaten the robustness of few-shot learning (FSL) due to the inexact features in a new domain. CLIP, a large-scale vision-language model, performs well in FSL on image-text embedding similarities, but it is…

Domain GeneralizationFew-Shot Learningzero-shot-classificationZero-Shot Learning

Self-Calibrated Consistency can Fight Back for Adversarial Robustness in Vision-Language Models

2025-10-26 · Jiaxiang Liu, Jiawei Du, Xiao Liu, Prayag Tiwari 외 arxiv

Pre-trained vision-language models (VLMs) such as CLIP have demonstrated strong zero-shot capabilities across diverse domains, yet remain highly vulnerable to adversarial perturbations that disrupt image-text alignment a…

Adversarial Robustness