paper-with-me

홈 › Papers

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

2026-08-13 · Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang arxiv

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

📄 PDF Abstract BibTeX arXiv:2608.13344

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Probing the Limits of Stylistic Alignment in Vision-Language Models

2025-09-29 · Asma Farajidizaji, Akash Gupta, Vatsal Raina arxiv

Vision-language models are increasingly used to generate image captions in specific styles, such as humor or romantic. However, these transformer-based models often struggle with this subjective task in a zero-shot setti…

MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly

2025-05-15 · Zhaowei Wang, Wenhao Yu, Xiyu Ren, Jipeng Zhang 외

The rapid extension of context windows in large vision-language models has given rise to long-context vision-language models (LCVLMs), which are capable of handling hundreds of images with interleaved text tokens in a si…

8kBenchmarkingRAG

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents

2025-06-02 · Bobo Li, Yuheng Wang, Hao Fei, Juncheng Li 외

Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one click", existing tools remain largely ru…

BenchmarkingForm

ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images

2025-11-30 · Yunfei Zhang, Yizhuo He, Yuanxun Shao, Zhengtao Yao 외 arxiv

Vision-Language Models (VLMs) have advanced multimodal understanding, yet still struggle when targets are embedded in cluttered backgrounds requiring figure-ground segregation. To address this, we introduce ChromouVQA, a…

Spatial Reasoning

Grounding Descriptions in Images informs Zero-Shot Visual Recognition

2024-12-05 · Shaunak Halbe, Junjiao Tian, K J Joseph, James Seale Smith 외

Vision-language models (VLMs) like CLIP have been cherished for their ability to perform zero-shot visual recognition on open-vocabulary concepts. This is achieved by selecting the object category whose textual represent…

AttributeBenchmarkingimage-classificationImage Classification+1