paper-with-me

홈 › Papers

MarineEval: Assessing the Marine Intelligence of Vision-Language Models

2025-12-24 · YuK-Kwan Wong, Tuan-An To, Jipeng Zhang, Ziqiang Zheng, Sai-Kit Yeung arxiv

We have witnessed promising progress led by large language models (LLMs) and further vision language models (VLMs) in handling various queries as a general-purpose assistant. VLMs, as a bridge to connect the visual world and language corpus, receive both visual content and various text-only user instructions to generate corresponding responses. Though great success has been achieved by VLMs in various fields, in this work, we ask whether the existing VLMs can act as domain experts, accurately answering marine questions, which require significant domain expertise and address special domain challenges/requirements. To comprehensively evaluate the effectiveness and explore the boundary of existing VLMs, we construct the first large-scale marine VLM dataset and benchmark called MarineEval, with 2,000 image-based question-answering pairs. During our dataset construction, we ensure the diversity and coverage of the constructed data: 7 task dimensions and 20 capacity dimensions. The domain requirements are specially integrated into the data construction and further verified by the corresponding marine domain experts. We comprehensively benchmark 17 existing VLMs on our MarineEval and also investigate the limitations of existing models in answering marine research questions. The experimental results reveal that existing VLMs cannot effectively answer the domain-specific questions, and there is still a large room for further performance improvements. We hope our new benchmark and observations will facilitate future research. Project Page: http://marineeval.hkustvgd.com/

📄 PDF Abstract BibTeX arXiv:2512.21126

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study

2024-01-04 · Ziqiang Zheng, YiWei Chen, Jipeng Zhang, Tuan-Anh Vu 외

Large language models (LLMs) have demonstrated a powerful ability to answer various queries as a general-purpose assistant. The continuous multi-modal large language models (MLLM) empower LLMs with the ability to perceiv…

Benchmarking Large Language Models for Image Classification of Marine Mammals

2024-10-22 · Yijiashun Qi, Shuzhang Cai, Zunduo Zhao, Jiaming Li 외

As Artificial Intelligence (AI) has developed rapidly over the past few decades, the new generation of AI, Large Language Models (LLMs) trained on massive datasets, has achieved ground-breaking performance in many applic…

Benchmarkingimage-classificationImage Classification

MARIDA: A benchmark for Marine Debris detection from Sentinel-2 remote sensing data

2022-01-07 · Plos one journal 2022 1 · Katerina Kikaki, Ioannis Kakogeorgiou, Paraskevi Mikeli, Dionysios E. Raitsos 외

Currently, a significant amount of research is focused on detecting Marine Debris and assessing its spectral behaviour via remote sensing, ultimately aiming at new operational monitoring solutions. Here, we introduce a M…

Image SegmentationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONSemantic Segmentation+3

Designing A Sustainable Marine Debris Clean-up Framework without Human Labels

2024-05-23 · Raymond Wang, Nicholas R. Record, D. Whitney King, Tahiya Chowdhury

Marine debris poses a significant ecological threat to birds, fish, and other animal life. Traditional methods for assessing debris accumulation involve labor-intensive and costly manual surveys. This study introduces a …

ClassificationLanguage ModellingObjectobject-detection+1

MarineGPT: Unlocking Secrets of Ocean to the Public

2023-10-20 · Ziqiang Zheng, Jipeng Zhang, Tuan-Anh Vu, Shizhe Diao 외

Large language models (LLMs), such as ChatGPT/GPT-4, have proven to be powerful tools in promoting the user experience as an AI assistant. The continuous works are proposing multi-modal large language models (MLLM), empo…

Language Modelling