paper-with-me

홈 › Papers

CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

2026-03-24 · Yuchen Wu, Kun Wang, Yining Pan, Na Zhao arxiv

Multi-modal fusion has emerged as a promising paradigm for accurate 3D object detection. However, performance degrades substantially when deployed in target domains different from training. In this work, focusing on dual-branch proposal-level detectors, we identify two factors that limit robust cross-domain generalization: 1) in challenging domains such as rain or nighttime, one modality may undergo severe degradation; 2) the LiDAR branch often dominates the detection process, leading to systematic underutilization of visual cues and vulnerability when point clouds are compromised. To address these challenges, we propose three components. First, Query-Decoupled Loss provides independent supervision for 2D-only, 3D-only, and fused queries, rebalancing gradient flow across modalities. Second, LiDAR-Guided Depth Prior augments 2D queries with instance-aware geometric priors through probabilistic fusion of image-predicted and LiDAR-derived depth distributions, improving their spatial initialization. Third, Complementary Cross-Modal Masking applies complementary spatial masks to the image and point cloud, encouraging queries from both modalities to compete within the fused decoder and thereby promoting adaptive fusion. Extensive experiments demonstrate substantial gains over state-of-the-art baselines while preserving source-domain performance. Code and models are publicly available at https://github.com/IMPL-Lab/CCF.

📄 PDF Abstract BibTeX arXiv:2603.23276

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Generalization3D Object DetectionPoint Clouds

Similar Papers 제목 키워드 기반

Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image Fusion

2025-08-21 · Mengyu Wang, Zhenyu Liu, Kun Li, Yu Wang 외 arxiv

Multimodal Image Fusion (MMIF) aims to integrate complementary information from different imaging modalities to overcome the limitations of individual sensors. It enhances image quality and facilitates downstream applica…

Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding

2025-02-11 · Ziyao Wang, Muneeza Azmart, Ang Li, Raya Horesh 외

Large Language Models (LLMs) often excel in specific domains but fall short in others due to the limitations of their training. Thus, enabling LLMs to solve problems collaboratively by integrating their complementary kno…

A Gated Cross-domain Collaborative Network for Underwater Object Detection

2023-06-25 · Linhui Dai, Hong Liu, Pinhao Song, Mengyuan Liu

Underwater object detection (UOD) plays a significant role in aquaculture and marine environmental protection. Considering the challenges posed by low contrast and low-light conditions in underwater environments, several…

2D Object DetectionImage Enhancementobject-detectionObject Detection+1

Reasoning with Autoregressive-Diffusion Collaborative Thoughts

2026-02-02 · Mu Yuan, Liekang Zeng, Guoliang Xing, Lan Zhang 외 arxiv

Autoregressive and diffusion models represent two complementary generative paradigms. Autoregressive models excel at sequential planning and constraint composition, yet struggle with tasks that require explicit spatial o…

Question AnsweringSpatial Reasoning

Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception

2023-07-26 · ICCV 2023 1 · Kun Yang, Dingkang Yang, Jingyu Zhang, Mingcheng Li 외

Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However,…

3D Object DetectionAutonomous Vehiclesobject-detectionObject Detection