paper-with-me

Papers

DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding

2025-05-27 · Weihao Xuan, Junjue Wang, Heli Qi, Zihang Chen, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya

Multimodal large language models have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite comprises 15,063 high-resolution (1.0m) multi-temporal images spanning 42 megacities in the U.S. from 2005 to 2023, organized into two components: DVL-Bench and DVL-Instruct. The DVL-Bench includes seven urban understanding tasks, from fundamental change detection (pixel-level) to quantitative analyses (regional-level) and comprehensive urban narratives (scene-level), capturing diverse urban dynamics including expansion/transformation patterns, disaster assessment, and environmental challenges. We evaluate 17 state-of-the-art multimodal large language models and reveal their limitations in long-term temporal understanding and quantitative analysis. These challenges motivate the creation of DVL-Instruct, a specialized instruction-tuning dataset designed to enhance models' capabilities in multi-temporal Earth observation. Building upon this dataset, we develop DVLChat, a baseline model capable of both image-level question-answering and pixel-level segmentation, facilitating a comprehensive understanding of city dynamics through language interactions.

📄 PDF Abstract BibTeX arXiv:2505.21076

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingChange DetectionEarth ObservationQuestion Answering

Similar Papers 제목 키워드 기반

DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

2026-01-29 · Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen 외 arxiv

Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models, which, despite strong generalization in static manipulation, struggle in dynamic scenarios requiring rapid perception, tempo…

Continuous Control

MMSpec: Benchmarking Speculative Decoding for Vision-Language Models

2026-03-16 · Hui Shen, Xin Wang, Ping Zhang, Yunta Hsieh 외 arxiv

Vision-language models (VLMs) achieve strong performance on multimodal tasks but suffer from high inference latency due to large model sizes and long multimodal contexts. Speculative decoding has recently emerged as an e…

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

2025-05-27 · Xuanwen Ding, Chengjun Pan, Zejun Li, Jiwen Zhang 외

Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introd…

BenchmarkingQuestion Selection

Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input

2025-10-19 · Chenxu Li, Zhicai Wang, Yuan Sheng, Xingyu Zhu 외 arxiv

Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robust…

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

2025-02-06 · He Hu, Yucheng Zhou, Lianzhong You, Hongbo Xu 외

With the integration of Multimodal large language models (MLLMs) into robotic systems and various AI applications, embedding emotional intelligence (EI) capabilities into these models is essential for enabling robots to …

BenchmarkingEmotional IntelligenceEmotion Recognition