paper-with-me

홈 › Papers

EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain

2024-01-30 · Wei zhang, Miaoxin Cai, Tong Zhang, Yin Zhuang, Xuerui Mao

Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversities between the natural and remote sensing (RS) images, the development of MLLMs in the RS domain is still in the infant stage. To fill the gap, a pioneer MLLM named EarthGPT integrating various multi-sensor RS interpretation tasks uniformly is proposed in this paper for universal RS image comprehension. In EarthGPT, three key techniques are developed including a visual-enhanced perception mechanism, a cross-modal mutual comprehension approach, and a unified instruction tuning method for multi-sensor multi-task in the RS domain. More importantly, a dataset named MMRS-1M featuring large-scale multi-sensor multi-modal RS instruction-following is constructed, comprising over 1M image-text pairs based on 34 existing diverse RS datasets and including multi-sensor images such as optical, synthetic aperture radar (SAR), and infrared. The MMRS-1M dataset addresses the drawback of MLLMs on RS expert knowledge and stimulates the development of MLLMs in the RS domain. Extensive experiments are conducted, demonstrating the EarthGPT's superior performance in various RS visual interpretation tasks compared with the other specialist models and MLLMs, proving the effectiveness of the proposed EarthGPT and offering a versatile paradigm for open-set reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2401.16822

Code (1)

wivizhang/earthgpt 공식 구현

Tasks

Image ComprehensionInstruction FollowingLanguage ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

EarthGPT-X: Enabling MLLMs to Flexibly and Comprehensively Understand Multi-Source Remote Sensing Imagery

2025-04-17 · Wei zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang 외

Recent advances in the visual-language area have developed natural multi-modal large language models (MLLMs) for spatial reasoning through visual prompting. However, due to remote sensing (RS) imagery containing abundant…

Large Language ModelMulti-Task LearningSpatial ReasoningVisual Prompting

E5-V: Universal Embeddings with Multimodal Large Language Models

2024-07-17 · Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang 외

Multimodal large language models (MLLMs) have shown promising advancements in general visual and language understanding. However, the representation of multimodal information using MLLMs remains largely unexplored. In th…

Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models

2025-06-02 · Youze Wang, WenBo Hu, Yinpeng Dong, Jing Liu 외

Large Language Models (LLMs) have evolved into Multimodal Large Language Models (MLLMs), significantly enhancing their capabilities by integrating visual information and other types, thus aligning more closely with the n…

Safety Alignment

Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models

2026-01-15 · Peng-Fei Zhang, Zi Huang arxiv

Existing adversarial attacks for VLP models are mostly sample-specific, resulting in substantial computational overhead when scaled to large datasets or new scenarios. To overcome this limitation, we propose Hierarchical…

AI2MMUM: AI-AI Oriented Multi-Modal Universal Model Leveraging Telecom Domain Large Model

2025-05-15 · Tianyu Jiao, Zhuoran Xiao, Yihang Huang, Chenhui Ye 외

Designing a 6G-oriented universal model capable of processing multi-modal data and executing diverse air interface tasks has emerged as a common goal in future wireless systems. Building on our prior work in communicatio…

Language ModelingLanguage ModellingLarge Language Modelmodel