paper-with-me

Papers

VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

2024-06-18 · Xiang Li, Jian Ding, Mohamed Elhoseiny

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this goal, existing datasets are typically tailored to single tasks, lack detailed object information, or suffer from inadequate quality control. Exploring these improvement opportunities, we present a Versatile vision-language Benchmark for Remote Sensing image understanding, termed VRSBench. This benchmark comprises 29,614 images, with 29,614 human-verified detailed captions, 52,472 object references, and 123,221 question-answer pairs. It facilitates the training and evaluation of vision-language models across a broad spectrum of remote sensing image understanding tasks. We further evaluated state-of-the-art models on this benchmark for three vision-language tasks: image captioning, visual grounding, and visual question answering. Our work aims to significantly contribute to the development of advanced vision-language models in the field of remote sensing. The data and code can be accessed at https://github.com/lx709/VRSBench.

📄 PDF Abstract BibTeX arXiv:2406.12384

Code (3)

lx709/vrsbench 공식 구현
lavender105/rsgpt pytorch
lx709/reobench pytorch

Tasks

Image CaptioningQuestion AnsweringVisual GroundingVisual Question Answering

Similar Papers 제목 키워드 기반

Visual Reasoning Agent: Robust Vision Systems in Remote Sensing via Inference-Time Scaling

2025-09-19 · Chung-En Johnny Yu, Brian Jalaian, Nathaniel D. Bastian arxiv

Building robust vision systems for high-stakes domains such as remote sensing requires stronger visual reasoning than what single-pass inference typically provides; yet, retraining large models is often computationally e…

Visual Reasoning

PaliGemma: A versatile 3B VLM for transfer

2024-07-10 · Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov 외

PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly knowledgeable base model that is effective t…

Language ModelingLanguage Modelling

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

2026-07-24 · Yuheng Zong, Minghua Wang, Xin Zhao, Zhi-Hui Zhan 외 arxiv

Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation applications require fine-grained scenario specialization, constrained by scarce hig…

VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation

2025-12-22 · Sihao Lin, Zerui Li, Xunyi Zhao, Gengze Zhou 외 arxiv

Despite remarkable progress in Vision-Language Navigation (VLN), existing benchmarks remain confined to fixed, small-scale datasets with naive physical simulation. These shortcomings limit the insight that the benchmarks…

Vision-Language Navigation

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

2024-07-03 · Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 외

We present InternLM-XComposer-2.5 (IXC-2.5), a versatile large-vision language model that supports long-contextual input and output. IXC-2.5 excels in various text-image comprehension and composition applications, achiev…

ArticlesImage ComprehensionLanguage ModelingLanguage Modelling+4