paper-with-me

홈 › Papers

SurvMamba: State Space Model with Multi-grained Multi-modal Interaction for Survival Prediction

2024-04-11 · Ying Chen, Jiajing Xie, Yuxiang Lin, Yuhang Song, Wenxian Yang, Rongshan Yu

Multi-modal learning that combines pathological images with genomic data has significantly enhanced the accuracy of survival prediction. Nevertheless, existing methods have not fully utilized the inherent hierarchical structure within both whole slide images (WSIs) and transcriptomic data, from which better intra-modal representations and inter-modal integration could be derived. Moreover, many existing studies attempt to improve multi-modal representations through attention mechanisms, which inevitably lead to high complexity when processing high-dimensional WSIs and transcriptomic data. Recently, a structured state space model named Mamba emerged as a promising approach for its superior performance in modeling long sequences with low complexity. In this study, we propose Mamba with multi-grained multi-modal interaction (SurvMamba) for survival prediction. SurvMamba is implemented with a Hierarchical Interaction Mamba (HIM) module that facilitates efficient intra-modal interactions at different granularities, thereby capturing more detailed local features as well as rich global representations. In addition, an Interaction Fusion Mamba (IFM) module is used for cascaded inter-modal interactive fusion, yielding more comprehensive features for survival prediction. Comprehensive evaluations on five TCGA datasets demonstrate that SurvMamba outperforms other existing methods in terms of performance and computational cost.

📄 PDF Abstract BibTeX arXiv:2404.08027

Code (0)

등록된 구현이 없습니다.

Tasks

MambaPredictionSurvival Predictionwhole slide images

Similar Papers 제목 키워드 기반

Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation

2025-11-14 · Quoc-Huy Trinh arxiv

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as…

Visual Grounding

MgSvF: Multi-Grained Slow vs. Fast Framework for Few-Shot Class-Incremental Learning

2020-06-28 · Hanbin Zhao, Yongjian Fu, Mintong Kang, Qi Tian 외

As a challenging problem, few-shot class-incremental learning (FSCIL) continually learns a sequence of tasks, confronting the dilemma between slow forgetting of old knowledge and fast adaptation to new knowledge. In this…

class-incremental learningClass Incremental LearningFew-Shot Class-Incremental LearningIncremental Learning

Video-Text Retrieval by Supervised Sparse Multi-Grained Learning

2023-02-19 · Yimu Wang, Peng Shi

While recent progress in video-text retrieval has been advanced by the exploration of better representation learning, in this paper, we present a novel multi-grained sparse learning framework, S3MA, to learn an aligned s…

Representation LearningRetrievalSparse LearningText Retrieval+2

F-HOI: Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions

2024-07-17 · Jie Yang, Xuesong Niu, Nan Jiang, Ruimao Zhang 외

Existing 3D human object interaction (HOI) datasets and models simply align global descriptions with the long HOI sequence, while lacking a detailed understanding of intermediate states and the transitions between states…

Human-Object Interaction DetectionLanguage ModellingLarge Language Model

Graph-based Fine-grained Multimodal Attention Mechanism for Sentiment Analysis

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multimodal sentiment analysis is a popular research area in natural language processing. Mainstream multimodal learning models barely consider that the visual and acoustic behaviors often have a much higher temporal freq…

multimodal interactionMultimodal Sentiment AnalysisSentiment Analysis