paper-with-me

홈 › Papers

LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency

2025-11-30 · Zhongbin Guo, Jiahe Liu, Wenyu Gao, Yushan Li, Chengzhi Li, Ping Jian arxiv

Text-driven 3D reconstruction requires masks that understand free-form instructions and remain stable under viewpoint changes. We present LISA-3D, a two-stage framework that adapts the instruction-following segmenter LISA with geometry-aware Low-Rank Adaptation (LoRA) layers while keeping the SAM-3D reconstructor frozen. During training, paired RGB-D frames and camera poses define a differentiable reprojection loss that enforces cross-view agreement without additional 3D-text annotations. At deployment, the adapted segmenter can produce an RGBA prompt for SAM-3D from one RGB image; when registered RGB-D views are available, optional logit fusion further improves the prompt. On ScanRefer and Nr3D, geometry-aware tuning improves both 2D masks and lifted 3D reconstructions while updating only 11.6M parameters. Our results separate geometry-aware training gains from optional multi-view inference gains, providing a modular route from language grounding to object-centric 3D reconstruction.

📄 PDF Abstract BibTeX arXiv:2512.01008

Code (0)

등록된 구현이 없습니다.

Tasks

Image Segmentation3D Reconstruction

Similar Papers 제목 키워드 기반

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

2025-05-05 · Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta 외

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent…

Reasoning SegmentationSegmentation

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

2024-09-29 · Zechen Bai, Tong He, Haiyang Mei, Pichao Wang 외

We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of l…

AllImage SegmentationLanguage ModelingLanguage Modelling+11

A Three-stage Approach for Segmenting Degraded Color Images: Smoothing, Lifting and Thresholding (SLaT)

2015-05-30 · Xiaohao Cai, Raymond Chan, Mila Nikolova, Tieyong Zeng

In this paper, we propose a SLaT (Smoothing, Lifting and Thresholding) method with three stages for multiphase segmentation of color images corrupted by different degradations: noise, information loss, and blur. At the f…

CPUSegmentation

Aerial Lifting: Neural Urban Semantic and Building Instance Lifting from Aerial Imagery

2024-03-18 · CVPR 2024 1 · Yuqi Zhang, GuanYing Chen, Jiaxing Chen, Shuguang Cui

We present a neural radiance field method for urban-scale semantic and building-level instance segmentation from aerial images by lifting noisy 2D labels to 3D. This is a challenging problem due to two primary reasons. F…

Instance SegmentationNeRFNovel View SynthesisSegmentation+1

LISA: Reasoning Segmentation via Large Language Model

2023-08-01 · CVPR 2024 1 · Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li 외

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition ta…

Language ModelingLanguage ModellingLarge Language Modelmodel+5