paper-with-me

Papers

Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling

2026-06-13 · Zhemin Zhang, Weijie Chen, David Le, Amara Tariq, Alex Wallace, Matthew Stib, Juan Maria Farina, Chadi Ayoub, Reza Arsanjani, Imon Banerjee arxiv

Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift. We introduce a foundation model-driven framework for cross-modal representation alignment between CT imaging and longitudinal EHR data, designed to generalize across tasks and institutions. CT and EHR modalities are encoded independently using domain-specific foundation models and aligned in a shared latent space through four principled fusion strategies: late fusion, contrastive alignment, cross-attention, and co-attention. We evaluate two clinically distinct TTE tasks: pulmonary embolism (PE) mortality and cardiovascular disease (CVD) outcomes, on large-scale multi-institutional cohorts (PE: N=3,099 train; 1,098 internal; 435 external; CVD: N=2,951 train; 837 internal; 682 external). Fusion consistently improves concordance index by 1.5-5.4% over unimodal baselines when modalities contribute comparably. Overall, contrastive multimodal fusion, particularly with CLMBR representations, provided the most consistent and statistically robust improvements, especially for PE mortality prediction. For MACE, cross-attention (one-hot) achieved the highest internal performance and image-guided co-attention achieved the best external performance. We therefore introduce a generalizable foundation model-based cross-modal alignment framework and provide the first systematic analysis of fusion behavior under modality imbalance in TTE prediction. Our results establish task-aware multimodal alignment as a necessary design principle for robust generalization and scalable clinical deployment.

📄 PDF Abstract BibTeX arXiv:2606.15038

Code (0)

등록된 구현이 없습니다.

Tasks

Mortality Prediction

Similar Papers 제목 키워드 기반

TMCIR: Token Merge Benefits Composed Image Retrieval

2025-04-15 · Chaoyang Wang, Zeyu Zhang, Long Teng, Zijun Li 외

Composed Image Retrieval (CIR) retrieves target images using a multi-modal query that combines a reference image with text describing desired modifications. The primary challenge is effectively fusing this visual and tex…

Contrastive Learningcross-modal alignmentImage RetrievalImage to text+1

SwimVG: Step-wise Multimodal Fusion and Adaption for Visual Grounding

2025-02-24 · Liangtao Shi, Ting Liu, Xiantao Hu, Yue Hu 외

Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-moda…

cross-modal alignmentVisual Grounding

Modality Fusion Network and Personalized Attention in Momentary Stress Detection in the Wild

2021-07-19 · Han Yu, Thomas Vaessen, Inez Myin-Germeys, Akane Sano

Multimodal wearable physiological data in daily life have been used to estimate self-reported stress labels. However, missing data modalities in data collection makes it challenging to leverage all the collected samples.…

Transfer Learning

Multimodal Fusion Learning with Dual Attention for Medical Imaging

2024-12-02 · Joy Dhar, Nayyar Zaidi, Maryam Haghighat, Puneet Goyal 외

Multimodal fusion learning has shown significant promise in classifying various diseases such as skin cancer and brain tumors. However, existing methods face three key limitations. First, they often lack generalizability…

User-Aware Conditional Generative Total Correlation Learning for Multi-Modal Recommendation

2026-04-03 · Jing Du, Zesheng Ye, Congbo Ma, Feng Liu 외 arxiv

Multi-modal recommendation (MMR) enriches item representations by introducing item content, e.g., visual and textual descriptions, to improve upon interaction-only recommenders. The success of MMR hinges on aligning thes…

Multi-modal Recommendation