paper-with-me

Papers

ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop

2026-03-03 · Shuangzhi Li, Lei Ma, Xingyu Li arxiv

Multi-modal 3D object detection is pivotal for autonomous driving, integrating complementary sensors like LiDAR and cameras. However, its real-world reliability is challenged by transient data interruptions and missing, where modalities can momentarily drop due to hardware glitches, adverse weather, or occlusions. This poses a critical risk, especially during a simultaneous modality drop, where the vehicle is momentarily blind. To address this problem, we introduce ModalPatch, the first plug-and-play module designed to enable robust detection under arbitrary modality-drop scenarios. Without requiring architectural changes or retraining, ModalPatch can be seamlessly integrated into diverse detection frameworks. Technically, ModalPatch leverages the temporal nature of sensor data for perceptual continuity, using a history-based module to predict and compensate for transiently unavailable features. To improve the fidelity of the predicted features, we further introduce an uncertainty-guided cross-modality fusion strategy that dynamically estimates the reliability of compensated features, suppressing biased signals while reinforcing informative ones. Extensive experiments show that ModalPatch consistently enhances both robustness and accuracy of state-of-the-art 3D object detectors under diverse modality-drop conditions.

📄 PDF Abstract BibTeX arXiv:2603.02481

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous Driving

Similar Papers 제목 키워드 기반

EvPlug: Learn a Plug-and-Play Module for Event and Image Fusion

2023-12-28 · Jianping Jiang, Xinyu Zhou, Peiqi Duan, Boxin Shi

Event cameras and RGB cameras exhibit complementary characteristics in imaging: the former possesses high dynamic range (HDR) and high temporal resolution, while the latter provides rich texture and color information. Th…

3D Hand Pose EstimationHand Pose Estimationobject-detectionObject Detection+2

Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization

2025-05-10 · Xu Zheng, Yuanhuiyi Lyu, Lutao Jiang, Danda Pani Paudel 외

Fusing and balancing multi-modal inputs from novel sensors for dense prediction tasks, particularly semantic segmentation, is critically important yet remains a significant challenge. One major limitation is the tendency…

SegmentationSemantic Segmentation

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

2026-02-26 · Siqi Lu, Wanying Xu, Yongbin Zheng, Wenting Luan 외 arxiv

Missing modalities present a fundamental challenge in multimodal models, often causing catastrophic performance degradation. Our observations suggest that this fragility stems from an imbalanced learning process, where t…

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

2023-04-27 · Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye 외

Large language models (LLMs) have demonstrated impressive zero-shot abilities on a variety of open-ended tasks, while recent research has also explored the use of LLMs for multi-modal generation. In this study, we introd…

Visual Question Answering (VQA)Zero-Shot Video Question Answer

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

2023-11-07 · CVPR 2024 1 · Qinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan 외

Multi-modal Large Language Models (MLLMs) have demonstrated impressive instruction abilities across various open-ended tasks. However, previous methods primarily focus on enhancing multi-modal capabilities. In this work,…

1 Image, 2*2 StitchingDecoderLanguage ModelingLanguage Modelling+4