paper-with-me

Papers

SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey

2024-10-31 · Kien X. Nguyen, Fengchun Qiao, Arthur Trembanis, Xi Peng

A major obstacle to the advancements of machine learning models in marine science, particularly in sonar imagery analysis, is the scarcity of AI-ready datasets. While there have been efforts to make AI-ready sonar image dataset publicly available, they suffer from limitations in terms of environment setting and scale. To bridge this gap, we introduce SeafloorAI, the first extensive AI-ready datasets for seafloor mapping across 5 geological layers that is curated in collaboration with marine scientists. We further extend the dataset to SeafloorGenAI by incorporating the language component in order to facilitate the development of both vision- and language-capable machine learning models for sonar imagery. The dataset consists of 62 geo-distributed data surveys spanning 17,300 square kilometers, with 696K sonar images, 827K annotated segmentation masks, 696K detailed language descriptions and approximately 7M question-answer pairs. By making our data processing source code publicly available, we aim to engage the marine science community to enrich the data pool and inspire the machine learning community to develop more robust models. This collaborative approach will enhance the capabilities and applications of our datasets within both fields.

📄 PDF Abstract BibTeX arXiv:2411.00172

Code (1)

deep-real/SeafloorAI 공식 구현

Similar Papers 제목 키워드 기반

BLIP3-KALE: Knowledge Augmented Large-Scale Dense Captions

2024-11-12 · Anas Awadalla, Le Xue, Manli Shu, An Yan 외

We introduce BLIP3-KALE, a dataset of 218 million image-text pairs that bridges the gap between descriptive synthetic captions and factual web-scale alt-text. KALE augments synthetic dense image captions with web-scale a…

DescriptiveImage Captioning

Utilizing Large Scale Vision and Text Datasets for Image Segmentation from Referring Expressions

2016-08-30 · Ronghang Hu, Marcus Rohrbach, Subhashini Venugopalan, Trevor Darrell

Image segmentation from referring expressions is a joint vision and language modeling task, where the input is an image and a textual expression describing a particular region in the image; and the goal is to localize an…

Image CaptioningImage SegmentationLanguage ModelingLanguage Modelling+2

VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment

2024-10-12 · Lei LI, Zhihui Xie, Mukai Li, Shunian Chen 외

As large vision-language models (LVLMs) evolve rapidly, the demand for high-quality and diverse data to align these models becomes increasingly crucial. However, the creation of such data with human supervision proves co…

DiversityHallucinationModels AlignmentRed Teaming+1

RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing

2023-06-20 · Zilun Zhang, Tiancheng Zhao, Yulong Guo, Jianwei Yin

Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A cr…

Cross-Modal RetrievalImage RetrievalImage-to-Text RetrievalLanguage Modeling+6

COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation

2025-09-10 · Umair Hassan arxiv

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems a…

Visual Grounding