paper-with-me

홈 › Papers

Language-Grounded Indoor 3D Semantic Segmentation in the Wild

2022-04-16 · David Rozenberszki, Or Litany, Angela Dai

Recent advances in 3D semantic segmentation with deep neural networks have shown remarkable success, with rapid performance increase on available datasets. However, current 3D semantic segmentation benchmarks contain only a small number of categories -- less than 30 for ScanNet and SemanticKITTI, for instance, which are not enough to reflect the diversity of real environments (e.g., semantic image understanding covers hundreds to thousands of classes). Thus, we propose to study a larger vocabulary for 3D semantic segmentation with a new extended benchmark on ScanNet data with 200 class categories, an order of magnitude more than previously studied. This large number of class categories also induces a large natural class imbalance, both of which are challenging for existing 3D semantic segmentation methods. To learn more robust 3D features in this context, we propose a language-driven pre-training method to encourage learned 3D features that might have limited training examples to lie close to their pre-trained text embeddings. Extensive experiments show that our approach consistently outperforms state-of-the-art 3D pre-training for 3D semantic segmentation on our proposed benchmark (+9% relative mIoU), including limited-data scenarios with +25% relative mIoU using only 5% annotations.

📄 PDF Abstract BibTeX arXiv:2204.07761

Code (1)

RozDavid/LanguageGroundedSemseg 공식 구현 pytorch

Tasks

3D Semantic SegmentationDiversitySegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

Towards Open-World Grasping with Large Vision-Language Models

2024-06-26 · Georgios Tziafas, Hamidreza Kasaei

The ability to grasp objects in-the-wild from open-ended language instructions constitutes a fundamental challenge in robotics. An open-world grasping system should be able to combine high-level contextual with low-level…

Robotic GraspingVisual GroundingVisual Prompting

A Self-Supervised Miniature One-Shot Texture Segmentation (MOSTS) Model for Real-Time Robot Navigation and Embedded Applications

2023-06-15 · Yu Chen, Chirag Rastogi, Zheyu Zhou, William R. Norris

Determining the drivable area, or free space segmentation, is critical for mobile robots to navigate indoor environments safely. However, the lack of coherent markings and structures (e.g., lanes, curbs, etc.) in indoor …

NavigateRobot NavigationSegmentationSemantic Segmentation

Open Panoramic Segmentation

2024-07-02 · Junwei Zheng, Ruiping Liu, Yufan Chen, Kunyu Peng 외

Panoramic images, capturing a 360{\deg} field of view (FoV), encompass omnidirectional spatial information crucial for scene understanding. However, it is not only costly to obtain training-sufficient dense-annotated pan…

Open-Vocabulary Panoramic Semantic Segmentation

FIRE-VLM: A Vision-Language-Driven Reinforcement Learning Framework for UAV Wildfire Tracking in a Physics-Grounded Fire Digital Twin

2026-01-06 · Chris Webb, Mobin Habibpour, Mayamin Hamid Raha, Ali Reza Tavakkoli 외 arxiv

Wildfire monitoring demands autonomous systems capable of reasoning under extreme visual degradation, rapidly evolving physical dynamics, and scarce real-world training data. Existing UAV navigation approaches rely on si…

Reinforcement Learning

JSMNet Improving Indoor Point Cloud Semantic and Instance Segmentation through Self-Attention and Multiscale

2023-09-14 · Shuochen Xu, Zhenxin Zhang

The semantic understanding of indoor 3D point cloud data is crucial for a range of subsequent applications, including indoor service robots, navigation systems, and digital twin engineering. Global features are crucial f…

Instance SegmentationSegmentationSemantic Segmentation