paper-with-me

Papers

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

2025-04-14 · Darryl Hannan, John Cooper, Dylan White, Timothy Doster, Henry Kvinge, Yijing Watkins

Multimodal large language models (MLLMs) have altered the landscape of computer vision, obtaining impressive results across a wide range of tasks, especially in zero-shot settings. Unfortunately, their strong performance does not always transfer to out-of-distribution domains, such as earth observation (EO) imagery. Prior work has demonstrated that MLLMs excel at some EO tasks, such as image captioning and scene understanding, while failing at tasks that require more fine-grained spatial reasoning, such as object localization. However, MLLMs are advancing rapidly and insights quickly become out-dated. In this work, we analyze more recent MLLMs that have been explicitly trained to include fine-grained spatial reasoning capabilities, benchmarking them on EO object localization tasks. We demonstrate that these models are performant in certain settings, making them well suited for zero-shot scenarios. Additionally, we provide a detailed discussion focused on prompt selection, ground sample distance (GSD) optimization, and analyzing failure cases. We hope that this work will prove valuable as others evaluate whether an MLLM is well suited for a given EO localization task and how to optimize it.

📄 PDF Abstract BibTeX arXiv:2504.10727

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingEarth ObservationImage CaptioningObject LocalizationScene UnderstandingSpatial Reasoning

Similar Papers 제목 키워드 기반

A Vision Centric Remote Sensing Benchmark

2025-03-20 · Abduljaleel Adejumo, Faegheh Yeganli, Clifford Broni-Bediako, Aoran Xiao 외

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike natural images, RS imagery presents uniqu…

Question AnsweringRepresentation LearningSpatial ReasoningVisual Grounding+2

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

2026-07-22 · Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang 외 arxiv

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery…

Visual Question AnsweringRelational ReasoningScene UnderstandingVisual Grounding

EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing

2025-03-30 · Hongxiang Jiang, Jihao Yin, Qixiong Wang, Jiaqi Feng 외

Recent advances in multimodal large language models (MLLMs) have demonstrated impressive results in various visual tasks. However, in remote sensing (RS), high resolution and small proportion of objects pose challenges t…

AttributeDisentanglementObjectobject-detection+1

ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG

2024-11-12 · Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Zian Guan 외

Ultra High Resolution (UHR) remote sensing imagery (RSI) (e.g. 100,000 $\times$ 100,000 pixels or more) poses a significant challenge for current Remote Sensing Multimodal Large Language Models (RSMLLMs). If choose to re…

RAGRetrievalRetrieval-augmented Generation

LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

2024-02-04 · Dilxat Muhtar, Zhenshi Li, Feng Gu, Xueliang Zhang 외

The revolutionary capabilities of large language models (LLMs) have paved the way for multimodal large language models (MLLMs) and fostered diverse applications across various specialized domains. In the remote sensing (…

Language ModelingLanguage Modelling