paper-with-me

홈 › Papers

Multi-view Vision-Prompt Fusion Network: Can 2D Pre-trained Model Boost 3D Point Cloud Data-scarce Learning?

2023-04-20 · Haoyang Peng, Baopu Li, Bo Zhang, Xin Chen, Tao Chen, Hongyuan Zhu

Point cloud based 3D deep model has wide applications in many applications such as autonomous driving, house robot, and so on. Inspired by the recent prompt learning in natural language processing, this work proposes a novel Multi-view Vision-Prompt Fusion Network (MvNet) for few-shot 3D point cloud classification. MvNet investigates the possibility of leveraging the off-the-shelf 2D pre-trained models to achieve the few-shot classification, which can alleviate the over-dependence issue of the existing baseline models towards the large-scale annotated 3D point cloud data. Specifically, MvNet first encodes a 3D point cloud into multi-view image features for a number of different views. Then, a novel multi-view prompt fusion module is developed to effectively fuse information from different views to bridge the gap between 3D point cloud data and 2D pre-trained models. A set of 2D image prompts can then be derived to better describe the suitable prior knowledge for a large-scale pre-trained image model for few-shot 3D point cloud classification. Extensive experiments on ModelNet, ScanObjectNN, and ShapeNet datasets demonstrate that MvNet achieves new state-of-the-art performance for 3D few-shot point cloud image classification. The source code of this work will be available soon.

📄 PDF Abstract BibTeX arXiv:2304.10224

Code (0)

등록된 구현이 없습니다.

Tasks

3D Point Cloud ClassificationAutonomous DrivingClassificationFew-Shot 3D Point Cloud Classificationimage-classificationImage ClassificationPoint Cloud ClassificationPrompt Learning

Similar Papers 제목 키워드 기반

Text-Image Conditioned Diffusion for Consistent Text-to-3D Generation

2023-12-19 · Yuze He, Yushi Bai, Matthieu Lin, Jenny Sheng 외

By lifting the pre-trained 2D diffusion models into Neural Radiance Fields (NeRFs), text-to-3D generation methods have made great progress. Many state-of-the-art approaches usually apply score distillation sampling (SDS)…

3D GenerationNeRFText to 3D

Multi-view Image Prompted Multi-view Diffusion for Improved 3D Generation

2024-04-26 · SeungWook Kim, Yichun Shi, Kejie Li, Minsu Cho 외

Using image as prompts for 3D generation demonstrate particularly strong performances compared to using text prompts alone, for images provide a more intuitive guidance for the 3D generation process. In this work, we del…

3D Generation

A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

2023-07-24 · Jindong Gu, Zhen Han, Shuo Chen, Ahmad Beirami 외

Prompt engineering is a technique that involves augmenting a large pre-trained model with task-specific hints, known as prompts, to adapt the model to new tasks. Prompts can be created manually as natural language instru…

Image GenerationImage-text matchingLanguage ModelingLanguage Modelling+5

In-Context Learning Unlocked for Diffusion Models

2023-05-01 · NeurIPS 2023 11 · Zhendong Wang, Yifan Jiang, Yadong Lu, Yelong Shen 외

We present Prompt Diffusion, a framework for enabling in-context learning in diffusion-based generative models. Given a pair of task-specific example images, such as depth from/to image and scribble from/to image, and a …

In-Context Learningtext-guided-image-editing

Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model

2024-04-28 · Xiaolong Li, Jiawei Mo, Ying Wang, Chethan Parameshwara 외

In this paper, we propose an effective two-stage approach named Grounded-Dreamer to generate 3D assets that can accurately follow complex, compositional text prompts while achieving high fidelity by using a pre-trained m…

Image GenerationText to 3D