paper-with-me

홈 › Papers

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

2026-06-25 · Minh Duc Do, Tillmann Rheude, Noel Kronenberg, Roland Eils, Benjamin Wild arxiv

Brain MRI poses a fundamental challenge for machine learning: models must learn from high-dimensional 3D data spanning multiple co-registered modalities, despite the limited sample sizes typical of neuroimaging studies relative to the diversity in anatomy, pathology, and acquisition conditions. While multimodal imaging provides complementary information critical for clinical interpretation, effectively integrating these signals remains difficult. We propose Multimodal Intra- and Cross-Context Vision Transformer (MICViT), a 3D vision transformer that explicitly models both modality-specific representations and cross-modal interactions across local and global contexts. Concretely, MICViT combines four attention mechanisms: modality-specific local and global attention for intra-modal feature learning, and cross-modal local and global attention to capture interactions between modalities. We evaluate MICViT on brain age prediction across three heterogeneous datasets (UK Biobank, n=41,404; SOOP, n=1,062; Cam-CAN, n=613) using multiple MRI modalities (e.g. T1, FLAIR, DWI, SWI). MICViT consistently outperforms state-of-the-art CNN and transformer baselines in 3D settings. Notably, it benefits more strongly from multimodal inputs, yielding larger performance gains as additional modalities are incorporated. These results demonstrate that explicitly modeling intra- and cross-modal interactions is key to unlocking the full potential of multimodal brain MRI, highlighting a promising direction for representation learning in neuroimaging.

📄 PDF Abstract BibTeX arXiv:2606.26894

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Global and Local Semantic Completion Learning for Vision-Language Pre-training

2023-06-12 · Rong-Cheng Tu, Yatai Ji, Jie Jiang, Weijie Kong 외

Cross-modal alignment plays a crucial role in vision-language pre-training (VLP) models, enabling them to capture meaningful associations across different modalities. For this purpose, numerous masked modeling tasks have…

cross-modal alignmentImage-text RetrievalLanguage ModellingMasked Language Modeling+5

Seeing What You Miss: Vision-Language Pre-training with Semantic Completion Learning

2022-11-24 · CVPR 2023 1 · Yatai Ji, RongCheng Tu, Jie Jiang, Weijie Kong 외

Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language mo…

cross-modal alignmentImage-text RetrievalLanguage ModelingLanguage Modelling+8

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

2020-05-01 · EMNLP 2020 11 · Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan 외

We present HERO, a novel framework for large-scale video+language omni-representation learning. HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-moda…

Language ModelingLanguage ModellingMasked Language ModelingMoment Retrieval+7

Centering Emotion Hotspots: Multimodal Local-Global Fusion and Cross-Modal Alignment for Emotion Recognition in Conversations

2025-10-07 · Yu Liu, Hanlei Shi, Haoxun Li, Yuqing Sun 외 arxiv

Emotion Recognition in Conversations (ERC) is hard because discriminative evidence is sparse, localized, and often asynchronous across modalities. We center ERC on emotion hotspots and present a unified model that detect…

Emotion Recognition

Mamba Goes HoME: Hierarchical Soft Mixture-of-Experts for 3D Medical Image Segmentation

2025-07-08 · Szymon Płotka, Gizem Mert, Maciej Chrabaszcz, Ewa Szczurek 외 arxiv

In recent years, artificial intelligence has significantly advanced medical image segmentation. Nonetheless, challenges remain, including efficient 3D medical image processing across diverse modalities and handling data …

Medical Image Segmentation