paper-with-me

홈 › Papers

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

2026-04-25 · Yida Xue, Ningyu Zhang, Tingwei Wu, Zhe Ma, Daxiong Ji, Zhao Wang, Guozhou Zheng, Huajun Chen arxiv

The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence has so far delivered limited impact in this domain due to a fundamental data bottleneck. Specifically, ocean data are highly fragmented across disparate sources and inherently exhibit multi-modal, high-noise, and weakly labeled characteristics, lacking unified schemas and semantic alignment. Although Multimodal Large Language Models (MLLMs) have achieved remarkable success in general domains, their application to ocean science remains severely constrained by the absence of large-scale, well-aligned multimodal datasets tailored to marine environments. To bridge this gap, we introduce OceanPile, a large-scale multimodal corpus designed for ocean foundation models. It comprises three key components: OceanCorpus, a unified collection integrating sonar data, underwater imagery, marine science visuals, and scientific text from diverse authoritative sources; OceanInstruction, a high-quality instruction dataset synthesized via a novel pipeline guided by a hierarchical Ocean Concept Knowledge Graph; and OceanBenchmark, a manually curated evaluation benchmark for rigorous assessment. We establish a multi-stage quality control process to ensure scientific validity and alignment across modalities. Experimental validation demonstrates significant performance improvements for models trained on our data. All datasets are publicly released to advance the field of marine artificial intelligence and empower domain-specific MLLMs.

📄 PDF Abstract BibTeX arXiv:2605.00877

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NIVA: A Multimodal Foundation Model for Actionable Earth System Intelligence

2026-06-26 · Anisha Pal, Aodhan Sweeney, Kyle Heyblom, Kalai Ramea arxiv

Recent advances in AI-driven weather and climate modeling have improved forecast skill while reducing computational cost. However, existing data-driven approaches are limited in their ability to model coupled Earth syste…

Fine-Tuning Large Multimodal Models for Automatic Pronunciation Assessment

2025-09-19 · Ke Wang, Wenning Wei, Yan Deng, Lei He 외 arxiv

Automatic Pronunciation Assessment (APA) is critical for Computer-Assisted Language Learning (CALL), requiring evaluation across multiple granularities and aspects. Large Multimodal Models (LMMs) present new opportunitie…

speechocean762: An Open-Source Non-native English Speech Corpus For Pronunciation Assessment

2021-04-03 · Junbo Zhang, Zhiwen Zhang, Yongqing Wang, Zhiyong Yan 외

This paper introduces a new open-source speech corpus named "speechocean762" designed for pronunciation assessment use, consisting of 5000 English utterances from 250 non-native speakers, where half of the speakers are c…

Phone-level pronunciation scoringSentencespeech-recognition

OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

2024-06-12 · Qingyun Li, Zhe Chen, Weiyun Wang, Wenhai Wang 외

Image-text interleaved data, consisting of multiple images and texts arranged in a natural document format, aligns with the presentation paradigm of internet data and closely resembles human reading habits. Recent studie…

In-Context Learning

Exploring the Underwater World Segmentation without Extra Training

2025-11-11 · Bingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao 외 arxiv

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce \t…