paper-with-me

Papers

SkyScript: A Large and Semantically Diverse Vision-Language Dataset for Remote Sensing

2023-12-20 · Zhecheng Wang, Rajanie Prabha, Tianyuan Huang, Jiajun Wu, Ram Rajagopal

Remote sensing imagery, despite its broad applications in helping achieve Sustainable Development Goals and tackle climate change, has not yet benefited from the recent advancements of versatile, task-agnostic vision language models (VLMs). A key reason is that the large-scale, semantically diverse image-text dataset required for developing VLMs is still absent for remote sensing images. Unlike natural images, remote sensing images and their associated text descriptions cannot be efficiently collected from the public Internet at scale. In this work, we bridge this gap by using geo-coordinates to automatically connect open, unlabeled remote sensing images with rich semantics covered in OpenStreetMap, and thus construct SkyScript, a comprehensive vision-language dataset for remote sensing images, comprising 2.6 million image-text pairs covering 29K distinct semantic tags. With continual pre-training on this dataset, we obtain a VLM that surpasses baseline models with a 6.2% average accuracy gain in zero-shot scene classification across seven benchmark datasets. It also demonstrates the ability of zero-shot transfer for fine-grained object attribute classification and cross-modal retrieval. We hope this dataset can support the advancement of VLMs for various multi-modal tasks in remote sensing, such as open-vocabulary classification, retrieval, captioning, and text-to-image synthesis.

📄 PDF Abstract BibTeX arXiv:2312.12856

Code (1)

wangzhecheng/skyscript 공식 구현 pytorch

Tasks

AttributeCross-Modal RetrievalImage GenerationRetrievalScene Classification

Similar Papers 제목 키워드 기반

SkyScript-100M: 1,000,000,000 Pairs of Scripts and Shooting Scripts for Short Drama

2024-08-18 · Jing Tang, Quanlu Jia, Yuqiang Xie, Zeyu Gong 외

Generating high-quality shooting scripts containing information such as scene and shot language is essential for short drama script generation. We collect 6,660 popular short drama episodes from the Internet, each with a…

Script GenerationVideo Captioning

GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing

2026-04-12 · Maram Hasan, Md Aminur Hossain, Savitra Roy, Souparna Bhowmik 외 arxiv

Effective foundation modeling in remote sensing requires spatially aligned heterogeneous modalities coupled with semantically grounded supervision, yet such resources remain limited at scale. We present GeoMeld, a large-…

Representation Learning

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

2025-01-01 · CVPR 2025 1 · Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan 외

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findi…

Semantic SimilaritySemantic Textual Similarity

Language Guided Skill Discovery

2024-06-07 · Seungeun Rho, Laura Smith, Tianyu Li, Sergey Levine 외

Skill discovery methods enable agents to learn diverse emergent behaviors without explicit rewards. To make learned skills useful for unknown downstream tasks, obtaining a semantically diverse repertoire of skills is ess…

Diversity

Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models

2026-03-28 · Yuhang Han, Yuyang Wu, Zhengbo Jiao, Yiyu Wang 외 arxiv

Reinforcement Learning from Verifiable Rewards (RLVR) has substantially enhanced the reasoning capabilities of large language models in abstract reasoning tasks. However, its application to Large Vision-Language Models (…

Reinforcement LearningMultimodal Reasoning