paper-with-me

홈 › Papers

UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification

2025-11-21 · Taixi Chen, Jingyun Chen, Nancy Guo arxiv

Inspired by the recent success of the Mamba architecture in vision and language domains, we introduce a Unified Attention-Mamba (UAM) backbone. Unlike previous hybrid approaches that integrate Attention and Mamba modules in fixed proportions, our unified design flexibly combines their capabilities within a single cohesive architecture, eliminating the need for manual ratio tuning and improving encode capability. We develop two UAM variants to comprehensively evaluate the benefits of this unified structure. Building on this backbone, we further propose a multimodal UAM framework that jointly performs cell-level classification and image segmentation. Experimental results demonstrate that UAM achieves state-of-the-art performance across both tasks on public benchmarks, surpassing leading image-based foundation models. It improves cell classification accuracy from 74\% to 78\% ($n$=349,882 cells), and tumor segmentation precision from 75\% to 80\% ($n$=406 patches).

📄 PDF Abstract BibTeX arXiv:2511.17355

Code (0)

등록된 구현이 없습니다.

Tasks

Tumor SegmentationImage Segmentation

Similar Papers 제목 키워드 기반

ML-Mamba: Efficient Multi-Modal Large Language Model Utilizing Mamba-2

2024-07-29 · Wenjun Huang, Jiakai Pan, Jiahao Tang, Yanyu Ding 외

Multimodal Large Language Models (MLLMs) have attracted much attention for their multifunctionality. However, traditional Transformer architectures incur significant overhead due to their secondary computational complexi…

Language ModelingLanguage ModellingLarge Language ModelMamba+1

Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking

2025-06-30 · Shiao Wang, Ju Huang, Qingchuan Ma, Jinfeng Gao 외

Combining traditional RGB cameras with bio-inspired event cameras for robust object tracking has garnered increasing attention in recent years. However, most existing multimodal tracking algorithms depend heavily on high…

MambaObject TrackingVisual Object Tracking

VL-Mamba: Exploring State Space Models for Multimodal Learning

2024-03-20 · Yanyuan Qiao, Zheng Yu, Longteng Guo, Sihan Chen 외

Multimodal large language models (MLLMs) have attracted widespread interest and have rich applications. However, the inherent attention mechanism in its Transformer structure requires quadratic complexity and results in …

Language ModelingLanguage ModellingLarge Language ModelMamba+3

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

2026-04-09 · Daichi Yashima, Shuhei Kurita, Yusuke Oda, Shuntaro Suzuki 외 arxiv

In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging because of their intricate temporal dependencies and substantial se…

Video Captioning

OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models

2025-03-11 · Jialv Zou, Bencheng Liao, Qian Zhang, Wenyu Liu 외

Recent advancements in unified multimodal understanding and visual generation (or multimodal generation) models have been hindered by their quadratic computational complexity and dependence on large-scale training data. …

GPUMambamultimodal generationState Space Models+1