paper-with-me

홈 › Papers

UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding

2025-10-21 · Da Zhang, Chenggang Rong, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao, Xuelong Li arxiv

Large vision-language models (VLMs) have achieved remarkable success in natural scene understanding, yet their application to underwater environments remains largely unexplored. Underwater imagery presents unique challenges including severe light attenuation, color distortion, and suspended particle scattering, while requiring specialized knowledge of marine ecosystems and organism taxonomy. To bridge this gap, we introduce UWBench, a comprehensive benchmark specifically designed for underwater vision-language understanding. UWBench comprises 15,003 high-resolution underwater images captured across diverse aquatic environments, encompassing oceans, coral reefs, and deep-sea habitats. Each image is enriched with human-verified annotations including 15,281 object referring expressions that precisely describe marine organisms and underwater structures, and 124,983 question-answer pairs covering diverse reasoning capabilities from object recognition to ecological relationship understanding. The dataset captures rich variations in visibility, lighting conditions, and water turbidity, providing a realistic testbed for model evaluation. Based on UWBench, we establish three comprehensive benchmarks: detailed image captioning for generating ecologically informed scene descriptions, visual grounding for precise localization of marine organisms, and visual question answering for multimodal reasoning about underwater environments. Extensive experiments on state-of-the-art VLMs demonstrate that underwater understanding remains challenging, with substantial room for improvement. Our benchmark provides essential resources for advancing vision-language research in underwater contexts and supporting applications in marine science, ecological monitoring, and autonomous underwater exploration. Our code and benchmark will be available.

📄 PDF Abstract BibTeX arXiv:2510.18262

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringMultimodal ReasoningScene UnderstandingObject Recognition

Similar Papers 제목 키워드 기반

Exploring the Underwater World Segmentation without Extra Training

2025-11-11 · Bingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao 외 arxiv

Accurate segmentation of marine organisms is vital for biodiversity monitoring and ecological assessment, yet existing datasets and models remain largely limited to terrestrial scenes. To bridge this gap, we introduce \t…

A Benchmark dataset for both underwater image enhancement and underwater object detection

2020-06-29 · Long Chen, Lei Tong, Feixiang Zhou, Zheheng Jiang 외

Underwater image enhancement is such an important vision task due to its significance in marine engineering and aquatic robot. It is usually work as a pre-processing step to improve the performance of high level vision t…

Image EnhancementImage Quality AssessmentObjectobject-detection+1

Visual enhancement and 3D representation for underwater scenes: a review

2025-05-03 · Guoxi Huang, Haoran Wang, Brett Seymour, Evan Kovacs 외

Underwater visual enhancement (UVE) and underwater 3D reconstruction pose significant challenges in computer vision and AI-based tasks due to complex imaging conditions in aquatic environments. Despite the development of…

3D Reconstruction

WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale Benchmark

2024-05-30 · Chunhui Zhang, Li Liu, Guanjie Huang, Hao Wen 외

Underwater object tracking (UOT) is a foundational task for identifying and tracing submerged entities in underwater video sequences. However, current UOT datasets suffer from limitations in scale, diversity of target ca…

Knowledge DistillationObject Tracking

Underwater Camouflaged Object Tracking Meets Vision-Language SAM2

2024-09-25 · Chunhui Zhang, Li Liu, Guanjie Huang, Zhipeng Zhang 외

Over the past decade, significant progress has been made in visual object tracking, largely due to the availability of large-scale datasets. However, these datasets have primarily focused on open-air scenarios and have l…

ObjectObject TrackingVideo SegmentationVideo Semantic Segmentation+1