paper-with-me

홈 › Papers

Unveiling the Depths: A Multi-Modal Fusion Framework for Challenging Scenarios

2024-02-19 · Jialei Xu, Xianming Liu, Junjun Jiang, Kui Jiang, Rui Li, Kai Cheng, Xiangyang Ji

Monocular depth estimation from RGB images plays a pivotal role in 3D vision. However, its accuracy can deteriorate in challenging environments such as nighttime or adverse weather conditions. While long-wave infrared cameras offer stable imaging in such challenging conditions, they are inherently low-resolution, lacking rich texture and semantics as delivered by the RGB image. Current methods focus solely on a single modality due to the difficulties to identify and integrate faithful depth cues from both sources. To address these issues, this paper presents a novel approach that identifies and integrates dominant cross-modality depth features with a learning-based framework. Concretely, we independently compute the coarse depth maps with separate networks by fully utilizing the individual depth cues from each modality. As the advantageous depth spreads across both modalities, we propose a novel confidence loss steering a confidence predictor network to yield a confidence map specifying latent potential depth areas. With the resulting confidence map, we propose a multi-modal fusion network that fuses the final depth in an end-to-end manner. Harnessing the proposed pipeline, our method demonstrates the ability of robust depth estimation in a variety of difficult scenarios. Experimental results on the challenging MS$^2$ and ViViD++ datasets demonstrate the effectiveness and robustness of our method.

📄 PDF Abstract BibTeX arXiv:2402.11826

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth Estimation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Robust RGB-D Fusion for Saliency Detection

2022-08-02 · Zongwei Wu, Shriarulmozhivarman Gobichettipalayam, Brahim Tamadazte, Guillaume Allibert 외

Efficiently exploiting multi-modal inputs for accurate RGB-D saliency detection is a topic of high interest. Most existing works leverage cross-modal interactions to fuse the two streams of RGB-D for intermediate feature…

Saliency Detection

Weakly supervised alignment and registration of MR-CT for cervical cancer radiotherapy

2024-05-21 · Jjahao Zhang, Yin Gu, Deyu Sun, Yuhua Gao 외

Cervical cancer is one of the leading causes of death in women, and brachytherapy is currently the primary treatment method. However, it is important to precisely define the extent of paracervical tissue invasion to impr…

Computed Tomography (CT)Image RegistrationOptical Flow Estimation

Unveiling the Potential of Text in High-Dimensional Time Series Forecasting

2025-01-13 · Xin Zhou, Weiqing Wang, Shilin Qu, Zhiqiang Zhang 외

Time series forecasting has traditionally focused on univariate and multivariate numerical data, often overlooking the benefits of incorporating multimodal information, particularly textual data. In this paper, we propos…

Time SeriesTime Series Forecasting

UniCat: Crafting a Stronger Fusion Baseline for Multimodal Re-Identification

2023-10-28 · Jennifer Crawford, Haoli Yin, Luke McDermott, Daniel Cummings

Multimodal Re-Identification (ReID) is a popular retrieval task that aims to re-identify objects across diverse data streams, prompting many researchers to integrate multiple modalities into a unified representation. Whi…

Retrieval

AIM: Adaptive Intra-Network Modulation for Balanced Multimodal Learning

2025-08-27 · Shu Shen, C. L. Philip Chen, Tong Zhang arxiv

Multimodal learning has significantly enhanced machine learning performance but still faces numerous challenges and limitations. Imbalanced multimodal learning is one of the problems extensively studied in recent works a…