paper-with-me

홈 › Papers

Enhancing Generalizability of Representation Learning for Data-Efficient 3D Scene Understanding

2024-06-17 · Yunsong Wang, Na Zhao, Gim Hee Lee

The field of self-supervised 3D representation learning has emerged as a promising solution to alleviate the challenge presented by the scarcity of extensive, well-annotated datasets. However, it continues to be hindered by the lack of diverse, large-scale, real-world 3D scene datasets for source data. To address this shortfall, we propose Generalizable Representation Learning (GRL), where we devise a generative Bayesian network to produce diverse synthetic scenes with real-world patterns, and conduct pre-training with a joint objective. By jointly learning a coarse-to-fine contrastive learning task and an occlusion-aware reconstruction task, the model is primed with transferable, geometry-informed representations. Post pre-training on synthetic data, the acquired knowledge of the model can be seamlessly transferred to two principal downstream tasks associated with 3D scene understanding, namely 3D object detection and 3D semantic segmentation, using real-world benchmark datasets. A thorough series of experiments robustly display our method's consistent superiority over existing state-of-the-art pre-training approaches.

📄 PDF Abstract BibTeX arXiv:2406.11283

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detection3D Semantic SegmentationContrastive Learningobject-detectionObject DetectionRepresentation LearningScene UnderstandingSemantic Segmentation

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Towards Robust Algorithms for Surgical Phase Recognition via Digital Twin-based Scene Representation

2024-10-26 · Hao Ding, Yuqian Zhang, Hongchao Shu, Xu Lian 외

Purpose: Surgical phase recognition (SPR) is an integral component of surgical data science, enabling high-level surgical analysis. End-to-end trained neural networks that predict surgical phase directly from videos have…

InformativenessScene UnderstandingSurgical phase recognition

In-Place Panoptic Radiance Field Segmentation with Perceptual Prior for 3D Scene Understanding

2024-10-06 · Shenghao Li

Accurate 3D scene representation and panoptic understanding are essential for applications such as virtual reality, robotics, and autonomous driving. However, challenges persist with existing methods, including precise 2…

2D Panoptic SegmentationAutonomous DrivingPanoptic SegmentationScene Understanding

Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

2026-04-10 · Tsuheng Hsu, Guiyu Liu, Juho Kannala, Janne Heikkilä arxiv

Recent works on 3D scene understanding leverage 2D masks from visual foundation models (VFMs) to supervise radiance fields, enabling instance-level 3D segmentation. However, the supervision signals from foundation models…

Representation LearningScene Understanding

CaesarNeRF: Calibrated Semantic Representation for Few-shot Generalizable Neural Rendering

2023-11-27 · Haidong Zhu, Tianyu Ding, Tianyi Chen, Ilya Zharkov 외

Generalizability and few-shot learning are key challenges in Neural Radiance Fields (NeRF), often due to the lack of a holistic understanding in pixel-level rendering. We introduce CaesarNeRF, an end-to-end approach that…

Few-Shot LearningNeRFNeural Rendering

A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding

2025-08-07 · Mahmoud Chick Zaouali, Todd Charter, Yehor Karpichev, Brandon Haworth 외 arxiv

Gaussian Splatting has rapidly emerged as a transformative technique for real-time 3D scene representation, offering a highly efficient and expressive alternative to Neural Radiance Fields (NeRF). Its ability to render c…

Scene Understanding