paper-with-me

Papers

SemanticSplat: Feed-Forward 3D Scene Understanding with Language-Aware Gaussian Fields

2025-06-11 · Qijing Li, Jingxiang Sun, Liang An, Zhaoqi Su, Hongwen Zhang, Yebin Liu

Holistic 3D scene understanding, which jointly models geometry, appearance, and semantics, is crucial for applications like augmented reality and robotic interaction. Existing feed-forward 3D scene understanding methods (e.g., LSM) are limited to extracting language-based semantics from scenes, failing to achieve holistic scene comprehension. Additionally, they suffer from low-quality geometry reconstruction and noisy artifacts. In contrast, per-scene optimization methods rely on dense input views, which reduces practicality and increases complexity during deployment. In this paper, we propose SemanticSplat, a feed-forward semantic-aware 3D reconstruction method, which unifies 3D Gaussians with latent semantic attributes for joint geometry-appearance-semantics modeling. To predict the semantic anisotropic Gaussians, SemanticSplat fuses diverse feature fields (e.g., LSeg, SAM) with a cost volume representation that stores cross-view feature similarities, enhancing coherent and accurate scene comprehension. Leveraging a two-stage distillation framework, SemanticSplat reconstructs a holistic multi-modal semantic feature field from sparse-view images. Experiments demonstrate the effectiveness of our method for 3D scene understanding tasks like promptable and open-vocabulary segmentation. Video results are available at https://semanticsplat.github.io.

📄 PDF Abstract BibTeX arXiv:2506.09565

Code (0)

등록된 구현이 없습니다.

Tasks

3D ReconstructionScene Understanding

Similar Papers 제목 키워드 기반

LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images

2026-05-22 · Yilong Liu, Wanhua Li, Chen Zhu-Tian, Hanspeter Pfister arxiv

We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language-aligned semantic features from sparse unposed mult…

Novel View SynthesisScene Understanding3D Reconstruction

EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding

2026-03-04 · Seungjun Lee, Zihan Wang, Yunsong Wang, Gim Hee Lee arxiv

Understanding a 3D scene immediately with its exploration is essential for embodied tasks, where an agent must construct and comprehend the 3D scene in an online and nearly real-time manner. In this study, we propose Emb…

Scene Understanding3D Reconstruction

GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

2026-08-18 · Qijian Tian, Zimeng Wu, Xuhong Wang, Lizhuang Ma 외 arxiv

Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations…

Referring Expression

Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

2025-07-08 · Aleksandar Jevtić, Christoph Reich, Felix Wimbauer, Oliver Hahn 외

Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC …

3D geometryDomain GeneralizationRepresentation LearningScene Understanding

VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

2026-07-02 · Marwane Hariat, David Filliat, Antoine Manzanera arxiv

Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojec…

3D Semantic SegmentationScene Understanding3D Reconstruction