Zero-Shot Transfer 3D Point Cloud Classification
3개 벤치마크 · 논문 11편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative Pretraining
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
Uni3D: Exploring Unified 3D Representation at Scale
PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning
PointCLIP: Point Cloud Understanding by CLIP
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
Papers
OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images
Recent open-world 3D representation learning methods using Vision-Language Models (VLMs) to align 3D point cloud with image-text information have shown superior 3D zero-shot performance. However, CAD-rendered images for …
Representation LearningTransfer LearningZero-shot 3D classificationZero-shot 3D Point Cloud Classification+3ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
This paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages. ShapeLLM is built upon…
3D geometry3D Object Captioning3D Point Cloud Classification3D Point Cloud Linear Classification+13Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training
Contrastive learning has emerged as a promising paradigm for 3D open-world understanding, i.e., aligning point cloud representation to image and text embedding space individually. In this paper, we introduce MixCon3D, a …
Contrastive LearningRetrievalText to 3DZero-shot 3D classification+1Uni3D: Exploring Unified 3D Representation at Scale
Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes…
3D Object ClassificationRetrievalZero-shot 3D classificationZero-shot 3D Point Cloud Classification+4ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
Though the success of CLIP-based training recipes in vision-language models, their scalability to more modalities (e.g., 3D, audio, etc.) is limited to large-scale data, which is expensive or even inapplicable for rare m…
3D ClassificationQuestion AnsweringRepresentation LearningTraining-free 3D Point Cloud Classification+2OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding
We introduce OpenShape, a method for learning multi-modal joint representations of text, image, and point clouds. We adopt the commonly used multi-modal contrastive learning framework for representation alignment, but wi…
3D Classification3D Shape RepresentationContrastive LearningImage Generation+3