paper-with-me

홈 › Papers

Towards Vision-Language Geo-Foundation Model: A Survey

2024-06-13 · Yue Zhou, Litong Feng, Yiping Ke, Xue Jiang, Junchi Yan, Xue Yang, Wayne Zhang

Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding. However, most methods rely on training with general image datasets, and the lack of geospatial data leads to poor performance on earth observation. Numerous geospatial image-text pair datasets and VLFMs fine-tuned on them have been proposed recently. These new approaches aim to leverage large-scale, multimodal geospatial data to build versatile intelligent models with diverse geo-perceptive capabilities, which we refer to as Vision-Language Geo-Foundation Models (VLGFMs). This paper thoroughly reviews VLGFMs, summarizing and analyzing recent developments in the field. In particular, we introduce the background and motivation behind the rise of VLGFMs, highlighting their unique research significance. Then, we systematically summarize the core technologies employed in VLGFMs, including data construction, model architectures, and applications of various multimodal geospatial tasks. Finally, we conclude with insights, issues, and discussions regarding future research directions. To the best of our knowledge, this is the first comprehensive literature review of VLGFMs. We keep tracing related works at https://github.com/zytx121/Awesome-VLGFM.

📄 PDF Abstract BibTeX arXiv:2406.09385

Code (1)

zytx121/awesome-vlgfm 공식 구현 pytorch

Tasks

Earth ObservationImage CaptioningImage-text RetrievalmodelQuestion AnsweringSurveyText RetrievalVisual GroundingVisual Question Answering

Similar Papers 제목 키워드 기반

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

2024-10-29 · Shenghao Xie, Wenqiang Zu, Mingyang Zhao, Duo Su 외

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to v…

Survey

Multi-Modal Foundation Models for Computational Pathology: A Survey

2025-03-12 · Dong Li, Guihong Wan, Xintao Wu, Xinyu Wu 외

Foundation models have emerged as a powerful paradigm in computational pathology (CPath), enabling scalable and generalizable analysis of histopathological images. While early developments centered on uni-modal models tr…

Surveywhole slide images

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

2023-09-18 · Chunyuan Li, Zhe Gan, Zhengyuan Yang, Jianwei Yang 외

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to gene…

Image GenerationSurveyText to Image GenerationText-to-Image Generation

Vision-and-Language Navigation Today and Tomorrow: A Survey in the Era of Foundation Models

2024-07-09 · Yue Zhang, Ziqiao Ma, Jialu Li, Yanyuan Qiao 외

Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development. The remarkable achievements of foundation models have shaped the chall…

Vision and Language Navigation

A Survey for Foundation Models in Autonomous Driving

2024-02-02 · Haoxiang Gao, Zhongruo Wang, Yaqian Li, Kaiwen Long 외

The advent of foundation models has revolutionized the fields of natural language processing and computer vision, paving the way for their application in autonomous driving (AD). This survey presents a comprehensive revi…

3D Object DetectionAutonomous DrivingCode Generationobject-detection+3