paper-with-me

Papers

A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation

2025-07-08 · Yuxin Zhang, Jiahao Yang, Zhe Chen, Wenjun Zhu, Jin Zhao, Yue Gao arxiv

Recently, large vision-language models (LVLMs) unleash powerful analysis capabilities for low Earth orbit (LEO) satellite Earth observation images in the data center. However, fast satellite motion, brief satellite-ground station (GS) contact windows, and large size of the images pose a data download challenge. To enable near real-time Earth observation applications (e.g., disaster and extreme weather monitoring), we should explore how to deploy LVLM in LEO satellite networks, and design SpaceVerse, an efficient satellite-ground synergistic LVLM inference system. To this end, firstly, we deploy compact LVLMs on satellites for lightweight tasks, whereas regular LVLMs operate on GSs to handle computationally intensive tasks. Then, we propose a computing and communication co-design framework comprised of a progressive confidence network and an attention-based multi-scale preprocessing, used to identify on-satellite inferring data, and reduce data redundancy before satellite-GS transmission, separately. We implement and evaluate SpaceVerse on real-world LEO satellite constellations and datasets, achieving a 31.2% average gain in accuracy and a 51.2% reduction in latency compared to state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2507.05731

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enabling Near-realtime Remote Sensing via Satellite-Ground Collaboration of Large Vision-Language Models

2025-10-28 · Zihan Li, Jiahao Yang, Yuxin Zhang, Zhe Chen 외 arxiv

Large vision-language models (LVLMs) have recently demonstrated great potential in remote sensing (RS) tasks (e.g., disaster monitoring) conducted by low Earth orbit (LEO) satellites. However, their deployment in real-wo…

Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

2023-12-12 · Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 외

We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for conn…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+5

Renormalized Connection for Scale-preferred Object Detection in Satellite Imagery

2024-09-09 · Fan Zhang, Lingling Li, Licheng Jiao, Xu Liu 외

Satellite imagery, due to its long-range imaging, brings with it a variety of scale-preferred tasks, such as the detection of tiny/small objects, making the precise localization and detection of small objects of interest…

object-detectionObject Detection

UAV-CodeAgents: Scalable UAV Mission Planning via Multi-Agent ReAct and Vision-Language Reasoning

2025-05-12 · Oleg Sautenkov, Yasheerah Yaqoot, Muhammad Ahsan Mustafa, Faryal Batool 외

We present UAV-CodeAgents, a scalable multi-agent framework for autonomous UAV mission generation, built on large language and vision-language models (LLMs/VLMs). The system leverages the ReAct (Reason + Act) paradigm to…

Fire Detection

Are VLMs Lost Between Sky and Space? LinkS$^2$Bench for UAV-Satellite Dynamic Cross-View Spatial Intelligence

2026-04-02 · Dian Liu, Jie Feng, Di Li, Yuhui Zheng 외 arxiv

Synergistic spatial intelligence between UAVs and satellites is indispensable for emergency response and security operations, as it uniquely integrates macro-scale global coverage with dynamic, real-time local perception…

Spatial Reasoning