paper-with-me

홈 › Papers

SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

2025-04-13 · Kaiyu Li, Zepeng Xin, Li Pang, Chao Pang, Yupeng Deng, Jing Yao, GuiSong Xia, Deyu Meng, Zhi Wang, Xiangyong Cao

Remote sensing has become critical for understanding environmental dynamics, urban planning, and disaster management. However, traditional remote sensing workflows often rely on explicit segmentation or detection methods, which struggle to handle complex, implicit queries that require reasoning over spatial context, domain knowledge, and implicit user intent. Motivated by this, we introduce a new task, \ie, geospatial pixel reasoning, which allows implicit querying and reasoning and generates the mask of the target region. To advance this task, we construct and release the first large-scale benchmark dataset called EarthReason, which comprises 5,434 manually annotated image masks with over 30,000 implicit question-answer pairs. Moreover, we propose SegEarth-R1, a simple yet effective language-guided segmentation baseline that integrates a hierarchical visual encoder, a large language model (LLM) for instruction parsing, and a tailored mask generator for spatial correlation. The design of SegEarth-R1 incorporates domain-specific adaptations, including aggressive visual token compression to handle ultra-high-resolution remote sensing images, a description projection module to fuse language and multi-scale features, and a streamlined mask prediction pipeline that directly queries description embeddings. Extensive experiments demonstrate that SegEarth-R1 achieves state-of-the-art performance on both reasoning and referring segmentation tasks, significantly outperforming traditional and LLM-based segmentation methods. Our data and code will be released at https://github.com/earth-insights/SegEarth-R1.

📄 PDF Abstract BibTeX arXiv:2504.09644

Code (1)

earth-insights/segearth-r1 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelSegmentation

Similar Papers 제목 키워드 기반

SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images

2025-12-23 · Zepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li 외 arxiv

Effectively grounding complex language to pixels in remote sensing (RS) images is a critical challenge for applications like disaster response and environmental monitoring. Current models can parse simple, single-target …

Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing

2025-12-22 · Xu Zhang, Junyao Ge, Yang Zheng, Kaitai Guo 외 arxiv

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-…

Reinforcement Learning

TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation

2026-03-19 · Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu 외 arxiv

Vision-language models (VLMs) have shown promise in earth observation (EO), yet they struggle with tasks that require grounding complex spatial reasoning in precise pixel-level visual representations. To address this pro…

Temporal SequencesSpatial ReasoningVisual Reasoning

GRASP: Geospatial pixel Reasoning viA Structured Policy learning

2025-08-23 · Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Jing Li 외 arxiv

Geospatial pixel reasoning aims to generate segmentation masks in remote sensing imagery directly from natural-language instructions. Most existing approaches follow a paradigm that fine-tunes multimodal large language m…

Reinforcement Learning

SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images

2025-12-09 · Kaiyu Li, Shengqi Zhang, Yujie Wang, Yupeng Deng 외 arxiv

Most existing methods for training-free open-vocabulary semantic segmentation are based on CLIP. While these approaches have made progress, they often face challenges in precise localization or require complex pipelines …

3D Semantic Segmentation2D Semantic SegmentationChange Detection