paper-with-me

Papers

Multimodal SAM-adapter for Semantic Segmentation

2025-09-12 · Iacopo Curti, Pierluigi Zama Ramirez, Alioscia Petrelli, Luigi Di Stefano arxiv

Semantic segmentation, a key task in computer vision with broad applications in autonomous driving, medical imaging, and robotics, has advanced substantially with deep learning. Nevertheless, current approaches remain vulnerable to challenging conditions such as poor lighting, occlusions, and adverse weather. To address these limitations, multimodal methods that integrate auxiliary sensor data (e.g., LiDAR, infrared) have recently emerged, providing complementary information that enhances robustness. In this work, we present MM SAM-adapter, a novel framework that extends the capabilities of the Segment Anything Model (SAM) for multimodal semantic segmentation. The proposed method employs an adapter network that injects fused multimodal features into SAM's rich RGB features. This design enables the model to retain the strong generalization ability of RGB features while selectively incorporating auxiliary modalities only when they contribute additional cues. As a result, MM SAM-adapter achieves a balanced and efficient use of multimodal information. We evaluate our approach on three challenging benchmarks, DeLiVER, FMB, and MUSES, where MM SAM-adapter delivers state-of-the-art performance. To further analyze modality contributions, we partition DeLiVER and FMB into RGB-easy and RGB-hard subsets. Results consistently demonstrate that our framework outperforms competing methods in both favorable and adverse conditions, highlighting the effectiveness of multimodal adaptation for robust scene understanding. The code is available at the following link: https://github.com/iacopo97/Multimodal-SAM-Adapter.

📄 PDF Abstract BibTeX arXiv:2509.10408

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationScene UnderstandingAutonomous Driving

Similar Papers 제목 키워드 기반

MANet: Fine-Tuning Segment Anything Model for Multimodal Remote Sensing Semantic Segmentation

2024-10-15 · Xianping Ma, Xiaokang Zhang, Man-on Pun, Bo Huang

Multimodal remote sensing data, collected from a variety of sensors, provide a comprehensive and integrated perspective of the Earth's surface. By employing multimodal fusion techniques, semantic segmentation offers more…

General KnowledgeSegmentationSemantic Segmentation

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

2024-08-02

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modal…

SegmentationSemantic SegmentationThermal Image Segmentation

Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning

2023-12-01 · Shaohua Dong, Yunhe Feng, Qing Yang, Yan Huang 외

Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch…

Decoderobject-detectionObject DetectionPrompt Learning+6

Parameter-Efficient Modality-Balanced Symmetric Fusion for Multimodal Remote Sensing Semantic Segmentation

2026-03-18 · Haocheng Li, Juepeng Zheng, Shuangxi Miao, Ruibo Lu 외 arxiv

Multimodal remote sensing semantic segmentation enhances scene interpretation by exploiting complementary physical cues from heterogeneous data. Although pretrained Vision Foundation Models (VFMs) provide strong general-…

Semantic Segmentation

ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer

2024-02-26 · Bowen Dong, Guanglei Yang, WangMeng Zuo, Lei Zhang

In this paper, we delve into the realm of vision transformers for continual semantic segmentation, a problem that has not been sufficiently explored in previous literature. Empirical investigations on the adaptation of e…

Continual Semantic SegmentationSegmentationSemantic Segmentation