paper-with-me

홈 › Papers

Image2Point: 3D Point-Cloud Understanding with 2D Image Pretrained Models

2021-06-08 · Chenfeng Xu, Shijia Yang, Tomer Galanti, Bichen Wu, Xiangyu Yue, Bohan Zhai, Wei Zhan, Peter Vajda, Kurt Keutzer, Masayoshi Tomizuka

3D point-clouds and 2D images are different visual representations of the physical world. While human vision can understand both representations, computer vision models designed for 2D image and 3D point-cloud understanding are quite different. Our paper explores the potential of transferring 2D model architectures and weights to understand 3D point-clouds, by empirically investigating the feasibility of the transfer, the benefits of the transfer, and shedding light on why the transfer works. We discover that we can indeed use the same architecture and pretrained weights of a neural net model to understand both images and point-clouds. Specifically, we transfer the image-pretrained model to a point-cloud model by copying or inflating the weights. We find that finetuning the transformed image-pretrained models (FIP) with minimal efforts -- only on input, output, and normalization layers -- can achieve competitive performance on 3D point-cloud classification, beating a wide range of point-cloud models that adopt task-specific architectures and use a variety of tricks. When finetuning the whole model, the performance improves even further. Meanwhile, FIP improves data efficiency, reaching up to 10.0 top-1 accuracy percent on few-shot classification. It also speeds up the training of point-cloud models by up to 11.1x for a target accuracy (e.g., 90 % accuracy). Lastly, we provide an explanation of the image to point-cloud transfer from the aspect of neural collapse. The code is available at: \url{https://github.com/chenfengxu714/image2point}.

📄 PDF Abstract BibTeX arXiv:2106.04180

Code (1)

chenfengxu714/image2point 공식 구현 pytorch

Tasks

3D Point Cloud ClassificationPoint Cloud ClassificationScene Segmentation

Similar Papers 제목 키워드 기반

CLIP-based Point Cloud Classification via Point Cloud to Image Translation

2024-08-07 · Shuvozit Ghose, Manyi Li, Yiming Qian, Yang Wang

Point cloud understanding is an inherently challenging problem because of the sparse and unordered structure of the point cloud in the 3D space. Recently, Contrastive Vision-Language Pre-training (CLIP) based point cloud…

ClassificationPoint Cloud ClassificationTranslation

CrossPoint: Self-Supervised Cross-Modal Contrastive Learning for 3D Point Cloud Understanding

2022-03-01 · CVPR 2022 1 · Mohamed Afham, Isuru Dissanayake, Dinithi Dissanayake, Amaya Dharmasiri 외

Manual annotation of large-scale point cloud dataset for varying tasks such as 3D object classification, segmentation and detection is often laborious owing to the irregular structure of point clouds. Self-supervised lea…

3D Object Classification3D Point Cloud Linear ClassificationContrastive LearningFew-Shot 3D Point Cloud Classification+1

ViPFormer: Efficient Vision-and-Pointcloud Transformer for Unsupervised Pointcloud Understanding

2023-03-25 · Hongyu Sun, Yongcai Wang, Xudong Cai, Xuewei Bai 외

Recently, a growing number of work design unsupervised paradigms for point cloud processing to alleviate the limitation of expensive manual annotation and poor transferability of supervised methods. Among them, CrossPoin…

3D Shape ClassificationContrastive LearningSemantic Segmentation

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

2025-07-17 · Yifan Xu, Chao Zhang, Hanqi Jiang, Xiaoyan Wang 외

Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend Large Language Models (LLMs) for tackli…

3D Point Cloud ReconstructionPoint cloud reconstructionScene Understanding

Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

2023-09-01 · Ziyu Guo, Renrui Zhang, Xiangyang Zhu, Yiwen Tang 외

We introduce Point-Bind, a 3D multi-modality model aligning point clouds with 2D image, language, audio, and video. Guided by ImageBind, we construct a joint embedding space between 3D and multi-modalities, enabling many…

3D Generation3D Question Answering (3D-QA)Generative 3D Object ClassificationInstruction Following+5