paper-with-me

홈 › Papers

Language-Assisted 3D Scene Understanding

2023-12-18 · Yanmin Wu, Qiankun Gao, Renrui Zhang, Jian Zhang

The scale and quality of point cloud datasets constrain the advancement of point cloud learning. Recently, with the development of multi-modal learning, the incorporation of domain-agnostic prior knowledge from other modalities, such as images and text, to assist in point cloud feature learning has been considered a promising avenue. Existing methods have demonstrated the effectiveness of multi-modal contrastive training and feature distillation on point clouds. However, challenges remain, including the requirement for paired triplet data, redundancy and ambiguity in supervised features, and the disruption of the original priors. In this paper, we propose a language-assisted approach to point cloud feature learning (LAST-PCL), enriching semantic concepts through LLMs-based text enrichment. We achieve de-redundancy and feature dimensionality reduction without compromising textual priors by statistical-based and training-free significant feature selection. Furthermore, we also delve into an in-depth analysis of the impact of text contrastive training on the point cloud. Extensive experiments validate that the proposed method learns semantically meaningful point cloud features and achieves state-of-the-art or comparable performance in 3D semantic segmentation, 3D object detection, and 3D scene classification tasks.

📄 PDF Abstract BibTeX arXiv:2312.11451

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detection3D Semantic SegmentationDimensionality Reductionfeature selectionobject-detectionObject DetectionScene ClassificationScene UnderstandingSemantic SegmentationTriplet

Similar Papers 제목 키워드 기반

Language-Assisted 3D Feature Learning for Semantic Scene Understanding

2022-11-25 · Junbo Zhang, Guofan Fan, Guanghan Wang, Zhengyuan Su 외

Learning descriptive 3D features is crucial for understanding 3D scenes with diverse objects and complex structures. However, it is usually unknown whether important geometric attributes and scene context obtain enough e…

DescriptiveInstance SegmentationObjectobject-detection+3

AS3D: 2D-Assisted Cross-Modal Understanding with Semantic-Spatial Scene Graphs for 3D Visual Grounding

2025-05-07 · Feng Xiao, Hongbin Xu, Guocan Zhao, Wenxiong Kang

3D visual grounding aims to localize the unique target described by natural languages in 3D scenes. The significant gap between 3D and language modalities makes it a notable challenge to distinguish multiple similar obje…

3D visual groundingGraph AttentionObjectRelational Reasoning+1

EndoChat: Grounded Multimodal Large Language Model for Endoscopic Surgery

2025-01-20 · Guankun Wang, Long Bai, Junyi Wang, Kun Yuan 외

Recently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

2025-08-02 · Dianyi Yang, Xihan Wang, Yu Gao, Shiyang Liu 외 arxiv

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing metho…

Scene Understanding

SEA-Vision: A Multilingual Benchmark for Comprehensive Document and Scene Text Understanding in Southeast Asia

2026-03-16 · Pengfei Yue, Xingran Zhao, Juntao Chen, Peng Hou 외 arxiv

Multilingual document and scene text understanding plays an important role in applications such as search, finance, and public services. However, most existing benchmarks focus on high-resource languages and fail to eval…

Visual Question AnsweringSpeaker VerificationLogical Reasoning