paper-with-me

홈 › Papers

OpenSU3D: Open World 3D Scene Understanding using Foundation Models

2024-07-19 · Rafay Mohiuddin, Sai Manoj Prakhya, Fiona Collins, Ziyuan Liu, André Borrmann

In this paper, we present a novel, scalable approach for constructing open set, instance-level 3D scene representations, advancing open world understanding of 3D environments. Existing methods require pre-constructed 3D scenes and face scalability issues due to per-point feature vector learning, limiting their efficacy with complex queries. Our method overcomes these limitations by incrementally building instance-level 3D scene representations using 2D foundation models, efficiently aggregating instance-level details such as masks, feature vectors, names, and captions. We introduce fusion schemes for feature vectors to enhance their contextual knowledge and performance on complex queries. Additionally, we explore large language models for robust automatic annotation and spatial reasoning tasks. We evaluate our proposed approach on multiple scenes from ScanNet and Replica datasets demonstrating zero-shot generalization capabilities, exceeding current state-of-the-art methods in open world 3D scene understanding.

📄 PDF Abstract BibTeX arXiv:2407.14279

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingSpatial ReasoningZero-shot Generalization

Similar Papers 제목 키워드 기반

OpenSUN3D: 1st Workshop Challenge on Open-Vocabulary 3D Scene Understanding

2024-02-23 · Francis Engelmann, Ayca Takmaz, Jonas Schult, Elisabetta Fedele 외

This report provides an overview of the challenge hosted at the OpenSUN3D Workshop on Open-Vocabulary 3D Scene Understanding held in conjunction with ICCV 2023. The goal of this workshop series is to provide a platform f…

Scene Understanding

Open Scene Understanding: Grounded Situation Recognition Meets Segment Anything for Helping People with Visual Impairments

2023-07-15 · Ruiping Liu, Jiaming Zhang, Kunyu Peng, Junwei Zheng 외

Grounded Situation Recognition (GSR) is capable of recognizing and interpreting visual scenes in a contextually intuitive way, yielding salient activities (verbs) and the involved entities (roles) depicted in images. In …

DecoderGrounded Situation RecognitionNavigateScene Understanding

Shading Annotations in the Wild

2017-05-02 · CVPR 2017 7 · Balazs Kovacs, Sean Bell, Noah Snavely, Kavita Bala

Understanding shading effects in images is critical for a variety of vision and graphics problems, including intrinsic image decomposition, shadow removal, image relighting, and inverse rendering. As is the case with oth…

Image RelightingIntrinsic Image DecompositionInverse RenderingShadow Removal

OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation

2025-12-09 · Yexin Liu, Manyuan Zhang, Yueze Wang, Hongyu Li 외 arxiv

Despite the promising progress in subject-driven image generation, current models often deviate from the reference identities and struggle in complex scenes with multiple subjects. To address this challenge, we introduce…

Image Generation

Deep Filter Banks for Texture Recognition and Segmentation

2015-06-01 · CVPR 2015 6 · Mircea Cimpoi, Subhransu Maji, Andrea Vedaldi

Research in texture recognition often concentrates on the problem of material recognition in uncluttered conditions, an assumption rarely met by applications. In this work we conduct a first study of material and describ…

Material RecognitionScene Recognition