OpenIns3D: Snap and Lookup for 3D Open-vocabulary Instance Segmentation
In this work, we introduce OpenIns3D, a new 3D-input-only framework for 3D open-vocabulary scene understanding. The OpenIns3D framework employs a "Mask-Snap-Lookup" scheme. The "Mask" module learns class-agnostic mask proposals in 3D point clouds, the "Snap" module generates synthetic scene-level images at multiple scales and leverages 2D vision-language models to extract interesting objects, and the "Lookup" module searches through the outcomes of "Snap" to assign category names to the proposed masks. This approach, yet simple, achieves state-of-the-art performance across a wide range of 3D open-vocabulary tasks, including recognition, object detection, and instance segmentation, on both indoor and outdoor datasets. Moreover, OpenIns3D facilitates effortless switching between different 2D detectors without requiring retraining. When integrated with powerful 2D open-world models, it achieves excellent results in scene understanding tasks. Furthermore, when combined with LLM-powered 2D models, OpenIns3D exhibits an impressive capability to comprehend and process highly complex text queries that demand intricate reasoning and real-world knowledge. Project page: https://zheninghuang.github.io/OpenIns3D/
Code (1)
Tasks
3D Open-Vocabulary Instance Segmentation3D Open-Vocabulary Object DetectionInstance Segmentationobject-detectionObject DetectionOpen Vocabulary Object DetectionScene UnderstandingSemantic SegmentationZero-shot 3D Point Cloud ClassificationSimilar Papers 제목 키워드 기반
OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
Understanding 3D scenes is pivotal for autonomous driving, robotics, and augmented reality. Recent semantic Gaussian Splatting approaches leverage large-scale 2D vision models to project 2D semantic features onto 3D scen…
Scene UnderstandingAutonomous DrivingOpenInst: A Simple Query-Based Method for Open-World Instance Segmentation
Open-world instance segmentation has recently gained significant popularitydue to its importance in many real-world applications, such as autonomous driving, robot perception, and remote sensing. However, previous method…
Autonomous DrivingInstance SegmentationOpen-World Instance SegmentationSegmentation+1Point2Graph: An End-to-end Point Cloud-based 3D Open-Vocabulary Scene Graph for Robot Navigation
Current open-vocabulary scene graph generation algorithms highly rely on both 3D scene point cloud data and posed RGB-D images and thus have limited applications in scenarios where RGB-D images or camera poses are not re…
3D Open-Vocabulary Object DetectionGraph GenerationObjectobject-detection+3Illustrative Language Understanding: Large-Scale Visual Grounding with Image Search
We introduce Picturebook, a large-scale lookup operation to ground language via {`}snapshots{'} of our physical world accessed through image search. For each word in a vocabulary, we extract the top-$k$ images from Googl…
General ClassificationImage RetrievalMachine TranslationNatural Language Inference+6Vocabulary-Free 3D Instance Segmentation with Vision and Language Assistant
Most recent 3D instance segmentation methods are open vocabulary, offering a greater flexibility than closed-vocabulary methods. Yet, they are limited to reasoning within a specific set of concepts, \ie the vocabulary, p…
3D Instance SegmentationInstance SegmentationSemantic Segmentation