paper-with-me

Papers

PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning

2024-03-13 · Qifeng Zhou, Wenliang Zhong, Yuzhi Guo, Michael Xiao, Hehuan Ma, Junzhou Huang

In the field of computational histopathology, both whole slide images (WSIs) and diagnostic captions provide valuable insights for making diagnostic decisions. However, aligning WSIs with diagnostic captions presents a significant challenge. This difficulty arises from two main factors: 1) Gigapixel WSIs are unsuitable for direct input into deep learning models, and the redundancy and correlation among the patches demand more attention; and 2) Authentic WSI diagnostic captions are extremely limited, making it difficult to train an effective model. To overcome these obstacles, we present PathM3, a multimodal, multi-task, multiple instance learning (MIL) framework for WSI classification and captioning. PathM3 adapts a query-based transformer to effectively align WSIs with diagnostic captions. Given that histopathology visual patterns are redundantly distributed across WSIs, we aggregate each patch feature with MIL method that considers the correlations among instances. Furthermore, our PathM3 overcomes data scarcity in WSI-level captions by leveraging limited WSI diagnostic caption data in the manner of multi-task joint learning. Extensive experiments with improved classification accuracy and caption generation demonstrate the effectiveness of our method on both WSI classification and captioning task.

📄 PDF Abstract BibTeX arXiv:2403.08967

Code (0)

등록된 구현이 없습니다.

Tasks

Caption GenerationDiagnosticimage-classificationImage ClassificationMultiple Instance Learningwhole slide images

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

PathMMU: A Massive Multimodal Expert-Level Benchmark for Understanding and Reasoning in Pathology

2024-01-29 · Yuxuan Sun, Hao Wu, Chenglu Zhu, Sunyi Zheng 외

The emergence of large multimodal models has unlocked remarkable potential in AI, particularly in pathology. However, the lack of specialized, high-quality benchmark impeded their development and precise evaluation. To a…

PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis

2025-08-28 · Ye Zhang, Yu Zhou, Jingwen Qi, Yongbing Zhang 외 arxiv

Deep learning based automated pathological diagnosis has markedly improved diagnostic efficiency and reduced variability between observers, yet its clinical adoption remains limited by opaque model decisions and a lack o…

Visual ReasoningText Generation

PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs

2026-03-10 · Jinyue Li, Yuci Liang, Qiankun Li, Xinheng Lyu 외 arxiv

Computational pathology demands both visual pattern recognition and dynamic integration of structured domain knowledge, including taxonomy, grading criteria, and clinical evidence. In practice, diagnostic reasoning requi…

PathMoE: Interpretable Multimodal Interaction Experts for Pediatric Brain Tumor Classification

2026-03-02 · Jian Yu, Joakim Nguyen, Jinrui Fang, Awais Naeem 외 arxiv

Accurate classification of pediatric central nervous system tumors remains challenging due to histological complexity and limited training data. While pathology foundation models have advanced whole-slide image (WSI) ana…

Brain Tumor Classification

PathMoG: A Pathway-Centric Modular Graph Neural Network for Multi-Omics Survival Prediction

2026-04-27 · Di Wang, Chupei Tang, Junxiao Kong, Jixiu Zhai 외 arxiv

Cancer survival prediction from multi-omics data remains challenging because prognostic signals are high-dimensional, heterogeneous, and distributed across interacting genes and pathways. We propose PathMoG, a pathway-ce…

Graph Neural Network