paper-with-me

Papers

Multi-modal Gated Mixture of Local-to-Global Experts for Dynamic Image Fusion

2023-02-02 · ICCV 2023 1 · Yiming Sun, Bing Cao, Pengfei Zhu, QinGhua Hu

Infrared and visible image fusion aims to integrate comprehensive information from multiple sources to achieve superior performances on various practical tasks, such as detection, over that of a single modality. However, most existing methods directly combined the texture details and object contrast of different modalities, ignoring the dynamic changes in reality, which diminishes the visible texture in good lighting conditions and the infrared contrast in low lighting conditions. To fill this gap, we propose a dynamic image fusion framework with a multi-modal gated mixture of local-to-global experts, termed MoE-Fusion, to dynamically extract effective and comprehensive information from the respective modalities. Our model consists of a Mixture of Local Experts (MoLE) and a Mixture of Global Experts (MoGE) guided by a multi-modal gate. The MoLE performs specialized learning of multi-modal local features, prompting the fused images to retain the local information in a sample-adaptive manner, while the MoGE focuses on the global information that complements the fused image with overall texture detail and contrast. Extensive experiments show that our MoE-Fusion outperforms state-of-the-art methods in preserving multi-modal image texture and contrast through the local-to-global dynamic learning paradigm, and also achieves superior performance on detection tasks. Our code will be available: https://github.com/SunYM2020/MoE-Fusion.

📄 PDF Abstract BibTeX arXiv:2302.01392

Code (1)

sunym2020/moe-fusion 공식 구현 pytorch

Tasks

Infrared And Visible Image Fusion

Similar Papers 제목 키워드 기반

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

2025-10-07 · Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun 외 arxiv

Emotion Recognition in Conversations (ERC) is hard because discriminative evidence is sparse, localized, and often asynchronous across modalities. We center ERC on emotion hotspots and present a unified model that detect…

Emotion Recognition

Conjugate Mixture Models for Clustering Multimodal Data

2020-12-09 · Vasil Khalidov, Florence Forbes, Radu Horaud

The problem of multimodal clustering arises whenever the data are gathered with several physically different sensors. Observations from different modalities are not necessarily aligned in the sense there there is no obvi…

Clusteringglobal-optimizationModel Selection

RoME: Role-aware Mixture-of-Expert Transformer for Text-to-Video Retrieval

2022-06-26 · Burak Satar, Hongyuan Zhu, Hanwang Zhang, Joo Hwee Lim

Seas of videos are uploaded daily with the popularity of social channels; thus, retrieving the most related video contents with user textual queries plays a more crucial role. Most methods consider only one joint embeddi…

Mixture-of-ExpertsRetrievalText to Video RetrievalVideo Retrieval

T2VLAD: Global-Local Sequence Alignment for Text-Video Retrieval

2021-04-20 · CVPR 2021 1 · Xiaohan Wang, Linchao Zhu, Yi Yang

Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions. The key to this problem is to measure text-video similarities in a joint embedding space. How…

RetrievalVideo Retrieval

Local Precise Refinement: A Dual-Gated Mixture-of-Experts for Enhancing Foundation Model Generalization against Spectral Shifts

2026-03-08 · Xi Chen, Maojun Zhang, Yu Liu, Shen Yan arxiv

Domain Generalization Semantic Segmentation (DGSS) in spectral remote sensing is severely challenged by spectral shifts across diverse acquisition conditions, which cause significant performance degradation for models de…

Domain GeneralizationSemantic Segmentation