paper-with-me

Papers

Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

2026-08-20 · Mohammad Arif Ul Alam arxiv

Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary information that is less affected by optical visibility, prior visual-sonar research has largely focused on feature alignment and nominal detection performance. We investigate cross-modal robustness as visual reliability deteriorates and assess whether pretrained visual foundation-model representations can be complemented by sonar under severe degradation. We use frozen DINOv2 as the visual encoder and construct a controlled five-level benchmark ranging from clean to extreme visual conditions. We compare conventional visual detection, frozen foundation-model representations, sonar context, fixed multimodal fusion, clean-trained adaptive gating, and degradation-aware gated fusion. Our method trains the fusion mechanism across the full range of degradation while keeping the visual and sonar encoders frozen, allowing modality contributions to adapt without fine-tuning the pretrained backbone. Under extreme combined degradation, the DINOv2 baseline achieves 0.4610 balanced accuracy, while degradation-aware visual-sonar fusion reaches 0.6152, a 33.5% relative improvement. The learned sonar contribution increases from 14.2% under clean conditions to 41.3% under extreme degradation, demonstrating adaptive redistribution of cross-modal reliance. Fusion provides the largest gains under severe turbidity and blur, whereas color attenuation alone yields little additional benefit. These results show that foundation-model representations remain valuable but insufficient under severe information loss, and that explicitly adapting fusion to modality reliability can improve robust underwater multimodal perception.

📄 PDF Abstract BibTeX arXiv:2608.19710

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 113
arxivsub/arXivSub_daily_arxiv ★ 4

Similar Papers 제목 키워드 기반

uScenes: A Multimodal RGB and 3D Sonar Dataset for Underwater Robot Perception

2026-08-28 · Trung Tien Dong, Zhenqi Wu, Aditya Penumarti, Zi-Hao Zhang 외 arxiv

Robust perception is essential for the deployment of autonomous underwater robots. However, optical cameras become unreliable under poor illumination and backscatter. Forward looking (2D) acoustic sensors remain effectiv…

Representation LearningScene UnderstandingObject DetectionPoint Clouds

A Sonar-Visual Dataset for Cross-Modal Underwater Robot Perception

2026-05-31 · Weitung Chen, Phil Tinn, Per Gunnar Auran, Martin Ludvigsen 외 arxiv

Underwater robots typically use both cameras and sonar for perception to leverage the rich semantic details of vision and the robust range measurements of acoustics. However, learning to map between these modalities via …

Deep Learning for Visual Navigation of Underwater Robots

2023-10-30 · M. Sunbeam

This paper aims to briefly survey deep learning methods for visual navigation of underwater robotics. The scope of this paper includes the visual perception of underwater robotics with deep learning methods, the availabl…

Deep LearningImitation Learningreinforcement-learningVisual Navigation

USIM and U0: A Vision-Language-Action Dataset and Model for General Underwater Robots

2025-10-09 · Junwen Gu, Zhiheng Wu, Pengxuan Si, Shuang Qiu 외 arxiv

Underwater environments pose unique challenges for robotic navigation and manipulation. While existing research has primarily focused on task-specific methods, studies on general-purpose intelligence for multi-task execu…

Pose Estimation

LU2Net: A Lightweight Network for Real-time Underwater Image Enhancement

2024-06-21 · Haodong Yang, Jisheng Xu, Zhiliang Lin, Jianping He

Computer vision techniques have empowered underwater robots to effectively undertake a multitude of tasks, including object tracking and path planning. However, underwater optical factors like light refraction and absorp…

Image EnhancementObject TrackingVideo Enhancement