paper-with-me

Papers

Falcon: A Remote Sensing Vision-Language Foundation Model

2025-03-14 · Kelu Yao, Nuo Xu, Rong Yang, Yingying Xu, Zhuoyan Gao, Titinunt Kitrungrotsakul, Yi Ren, Pu Zhang, Jin Wang, Ning Wei, Chao Li

This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensing tasks. Falcon demonstrates powerful understanding and reasoning abilities at the image, region, and pixel levels. Specifically, given simple natural language instructions and remote sensing images, Falcon can produce impressive results in text form across 14 distinct tasks, i.e., image classification, object detection, segmentation, image captioning, and etc. To facilitate Falcon's training and empower its representation capacity to encode rich spatial and semantic information, we developed Falcon_SFT, a large-scale, multi-task, instruction-tuning dataset in the field of remote sensing. The Falcon_SFT dataset consists of approximately 78 million high-quality data samples, covering 5.6 million multi-spatial resolution and multi-view remote sensing images with diverse instructions. It features hierarchical annotations and undergoes manual sampling verification to ensure high data quality and reliability. Extensive comparative experiments are conducted, which verify that Falcon achieves remarkable performance over 67 datasets and 14 tasks, despite having only 0.7B parameters. We release the complete dataset, code, and model weights at https://github.com/TianHuiLab/Falcon, hoping to help further develop the open-source community.

📄 PDF Abstract BibTeX arXiv:2503.11070

Code (1)

tianhuilab/falcon 공식 구현 pytorch

Tasks

Image Captioningimage-classificationImage Classificationobject-detectionObject Detection

Similar Papers 제목 키워드 기반

PIR: Remote Sensing Image-Text Retrieval with Prior Instruction Representation Learning

2024-05-16 · Jiancheng Pan, Muyuan Ma, Qing Ma, Cong Bai 외

Remote sensing image-text retrieval constitutes a foundational aspect of remote sensing interpretation tasks, facilitating the alignment of vision and language representations. This paper introduces a prior instruction r…

Image-text RetrievalRepresentation LearningRetrievalScene Recognition+1

Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment

2023-12-12 · Utkarsh Mall, Cheng Perng Phoo, Meilin Kelsey Liu, Carl Vondrick 외

We introduce a method to train vision-language models for remote-sensing images without using any textual annotations. Our key insight is to use co-located internet imagery taken on the ground as an intermediary for conn…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+5

FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models

2026-06-15 · Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang 외 arxiv

Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared …

Representation LearningText Retrieval

Rethinking Electro-Optical Vision Foundation Models for Remote Sensing Retrieval: A Controlled Comparison with Generalist VFM

2026-05-04 · Hyobin Park, Minseok Seo, Dong-Geol Choi arxiv

Vision foundation models have attracted significant attention for their ability to leverage large-scale unlabeled visual data. This advantage is particularly important in remote sensing, where data acquisition is costly …

Image Retrieval

RemoteCLIP: A Vision Language Foundation Model for Remote Sensing

2023-06-19 · Fan Liu, Delong Chen, Zhangqingyun Guan, Xiaocong Zhou 외

General-purpose foundation models have led to recent breakthroughs in artificial intelligence. In remote sensing, self-supervised learning (SSL) and Masked Image Modeling (MIM) have been adopted to build foundation model…

ClassificationCross-Modal Retrievalimage-classificationImage Classification+8