paper-with-me

Papers

A large-scale image-text dataset benchmark for farmland segmentation

2025-03-29 · Chao Tao, Dandan Zhong, Weiliang Mu, Zhuofei Du, Haiyang Wu

The traditional deep learning paradigm that solely relies on labeled data has limitations in representing the spatial relationships between farmland elements and the surrounding environment.It struggles to effectively model the dynamic temporal evolution and spatial heterogeneity of farmland. Language,as a structured knowledge carrier,can explicitly express the spatiotemporal characteristics of farmland, such as its shape, distribution,and surrounding environmental information.Therefore,a language-driven learning paradigm can effectively alleviate the challenges posed by the spatiotemporal heterogeneity of farmland.However,in the field of remote sensing imagery of farmland,there is currently no comprehensive benchmark dataset to support this research direction.To fill this gap,we introduced language based descriptions of farmland and developed FarmSeg-VL dataset,the first fine-grained image-text dataset designed for spatiotemporal farmland segmentation.Firstly, this article proposed a semi-automatic annotation method that can accurately assign caption to each image, ensuring high data quality and semantic richness while improving the efficiency of dataset construction.Secondly,the FarmSeg-VL exhibits significant spatiotemporal characteristics.In terms of the temporal dimension,it covers all four seasons.In terms of the spatial dimension,it covers eight typical agricultural regions across China.In addition, in terms of captions,FarmSeg-VL covers rich spatiotemporal characteristics of farmland,including its inherent properties,phenological characteristics, spatial distribution,topographic and geomorphic features,and the distribution of surrounding environments.Finally,we present a performance analysis of VLMs and the deep learning models that rely solely on labels trained on the FarmSeg-VL,demonstrating its potential as a standard benchmark for farmland segmentation.

📄 PDF Abstract BibTeX arXiv:2503.23106

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CCMB: A Large-scale Chinese Cross-modal Benchmark

2022-05-08 · Chunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong 외

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets a…

image-classificationImage ClassificationImage GenerationImage Retrieval+9

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

2022-02-14 · Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 외

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…

BenchmarkingContrastive Learningimage-classificationImage Classification+6

CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios

2024-01-19 · Xiangshuo Qiao, Xianxin Li, Xiaozhe Qu, Jie Zhang 외

Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-training are presented in the form of open dom…

Common Sense ReasoningImage Retrieval

Chinese Street View Text: Large-scale Chinese Text Reading with Partially Supervised Learning

2019-09-17 · ICCV 2019 10 · Yipeng Sun, Jiaming Liu, Wei Liu, Junyu Han 외

Most existing text reading benchmarks make it difficult to evaluate the performance of more advanced deep learning models in large vocabularies due to the limited amount of training data. To address this issue, we introd…

Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model

2025-01-01 · Chenyang Liu, Keyan Chen, Rui Zhao, Zhengxia Zou 외

Generative foundation models have advanced large-scale text-driven natural image generation, becoming a prominent research trend across various vertical domains. However, in the remote sensing field, there is still a lac…

Image Generation