paper-with-me

홈 › Papers

PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding

2026-04-17 · Junjie Wen, Junlin He, Fei Ma, Jinqiang Cui arxiv

Accurate open-vocabulary 3D scene understanding requires semantic representations that are both language-aligned and spatially precise at the pixel level, while remaining scalable when lifted to 3D space. However, existing representations struggle to jointly satisfy these requirements, and densely propagating pixel-wise semantics to 3D often results in substantial redundancy, leading to inefficient storage and querying in large-scale scenes. To address these challenges, we present \emph{PLAF}, a Pixel-wise Language-Aligned Feature extraction framework that enables dense and accurate semantic alignment in 2D without sacrificing open-vocabulary expressiveness. Building upon this representation, we further design an efficient semantic storage and querying scheme that significantly reduces redundancy across both 2D and 3D domains. Experimental results show that \emph{PLAF} provides a strong semantic foundation for accurate and efficient open-vocabulary 3D scene understanding. The codes are publicly available at https://github.com/RockWenJJ/PLAF.

📄 PDF Abstract BibTeX arXiv:2604.15770

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Reliable Multi-modal Medical Image-to-image Translation Independent of Pixel-wise Aligned Data

2024-08-26 · Langrui Zhou, Guang Li

The current mainstream multi-modal medical image-to-image translation methods face a contradiction. Supervised methods with outstanding performance rely on pixel-wise aligned training data to constrain the model optimiza…

Image RegistrationImage-to-Image TranslationMedical Image RegistrationModel Optimization+1

VPFusion: Joint 3D Volume and Pixel-Aligned Feature Fusion for Single and Multi-view 3D Reconstruction

2022-03-14 · Jisan Mahmud, Jan-Michael Frahm

We introduce a unified single and multi-view neural implicit 3D reconstruction framework VPFusion. VPFusion attains high-quality reconstruction using both - 3D feature volume to capture 3D-structure-aware context, and pi…

3D Reconstruction3D Shape ReconstructionMulti-View 3D Reconstruction

4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models

2025-03-13 · CVPR 2025 1 · Wanhua Li, Renping Zhou, Jiawei Zhou, Yingwei Song 외

Learning 4D language fields to enable time-sensitive, open-ended language queries in dynamic scenes is essential for many real-world applications. While LangSplat successfully grounds CLIP features into 3D Gaussian repre…

Large Language ModelObjectSentence Embeddings

A System for 3D Reconstruction Of Comminuted Tibial Plafond Bone Fractures

2021-02-23 · Pengcheng Liu, Nathan Hewitt, Waseem Shadid, Andrew Willis

High energy impacts at joint locations often generate highly fragmented, or comminuted, bone fractures. Current approaches for treatment require physicians to decide how to classify the fracture within a hierarchy fractu…

3D ReconstructionAnatomyPrognosis

CARIS: Context-Augmented Referring Image Segmentation

2023-10-27 · ACM MM 2023 10 · Sun-Ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie 외

Referring image segmentation aims to segment the target object described by a natural-language utterance. Recent approaches typically distinguish pixels by aligning pixel-wise visual features with linguistic features ext…

DecoderImage SegmentationSegmentationSemantic Segmentation