paper-with-me

홈 › Papers

LISA: Reasoning Segmentation via Large Language Model

2023-08-01 · CVPR 2024 1 · Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, Jiaya Jia

Although perception systems have made remarkable advancements in recent years, they still rely on explicit human instruction or pre-defined categories to identify the target objects before executing visual recognition tasks. Such systems cannot actively reason and comprehend implicit user intention. In this work, we propose a new segmentation task -- reasoning segmentation. The task is designed to output a segmentation mask given a complex and implicit query text. Furthermore, we establish a benchmark comprising over one thousand image-instruction-mask data samples, incorporating intricate reasoning and world knowledge for evaluation purposes. Finally, we present LISA: large Language Instructed Segmentation Assistant, which inherits the language generation capabilities of multimodal Large Language Models (LLMs) while also possessing the ability to produce segmentation masks. We expand the original vocabulary with a <SEG> token and propose the embedding-as-mask paradigm to unlock the segmentation capability. Remarkably, LISA can handle cases involving complex reasoning and world knowledge. Also, it demonstrates robust zero-shot capability when trained exclusively on reasoning-free datasets. In addition, fine-tuning the model with merely 239 reasoning segmentation data samples results in further performance enhancement. Both quantitative and qualitative experiments show our method effectively unlocks new reasoning segmentation capabilities for multimodal LLMs. Code, models, and data are available at https://github.com/dvlab-research/LISA.

📄 PDF Abstract BibTeX arXiv:2308.00692

Code (2)

dvlab-research/lisa 공식 구현 pytorch
sunsmarterjie/chatterbox pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelmodelReasoning SegmentationReferring Video Object SegmentationSegmentationText GenerationWorld Knowledge

Similar Papers 제목 키워드 기반

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

2024-09-29 · Zechen Bai, Tong He, Haiyang Mei, Pichao Wang 외

We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of l…

AllImage SegmentationLanguage ModelingLanguage Modelling+11

LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

2023-12-28 · Senqiao Yang, Tianyuan Qu, Xin Lai, Zhuotao Tian 외

While LISA effectively bridges the gap between segmentation and large language models to enable reasoning segmentation, it poses certain limitations: unable to distinguish different instances of the target region, and co…

Instance SegmentationLanguage ModelingLanguage ModellingLarge Language Model+3

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

2025-05-05 · Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta 외

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent…

Reasoning SegmentationSegmentation

Beyond Segmentation: Road Network Generation with Multi-Modal LLMs

2023-10-15 · Sumedh Rasal, Sanjay Kumar Boddhu

This paper introduces an innovative approach to road network generation through the utilization of a multi-modal Large Language Model (LLM). Our model is specifically designed to process aerial images of road layouts and…

Autonomous NavigationLanguage ModelingLanguage ModellingLarge Language Model+1

Empowering Segmentation Ability to Multi-modal Large Language Models

2024-03-21 · YuQi Yang, Peng-Tao Jiang, Jing Wang, Hao Zhang 외

Multi-modal large language models (MLLMs) can understand image-language prompts and demonstrate impressive reasoning ability. In this paper, we extend MLLMs' output by empowering MLLMs with the segmentation ability. The …

Dialogue GenerationReasoning SegmentationSegmentationWord Embeddings