paper-with-me

Papers

WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image

2024-12-03 · Yuci Liang, Xinheng Lyu, Meidan Ding, WenTing Chen, Jipeng Zhang, Yuexiang Ren, Xiangjian He, Song Wu, Sen yang, Xiyue Wang, Xiaohan Xing, Linlin Shen

Recent advancements in computational pathology have produced patch-level Multi-modal Large Language Models (MLLMs), but these models are limited by their inability to analyze whole slide images (WSIs) comprehensively and their tendency to bypass crucial morphological features that pathologists rely on for diagnosis. To address these challenges, we first introduce WSI-Bench, a large-scale morphology-aware benchmark containing 180k VQA pairs from 9,850 WSIs across 30 cancer types, designed to evaluate MLLMs' understanding of morphological characteristics crucial for accurate diagnosis. Building upon this benchmark, we present WSI-LLaVA, a novel framework for gigapixel WSI understanding that employs a three-stage training approach: WSI-text alignment, feature space alignment, and task-specific instruction tuning. To better assess model performance in pathological contexts, we develop two specialized WSI metrics: WSI-Precision and WSI-Relevance. Experimental results demonstrate that WSI-LLaVA outperforms existing models across all capability dimensions, with a significant improvement in morphological analysis, establishing a clear correlation between morphological understanding and diagnostic accuracy.

📄 PDF Abstract BibTeX arXiv:2412.02141

Code (0)

등록된 구현이 없습니다.

Tasks

DiagnosticLanguage ModelingLanguage ModellingLarge Language ModelMorphological AnalysisMultimodal Large Language ModelVisual Question Answering (VQA)whole slide images

Similar Papers 제목 키워드 기반

Efficient Whole Slide Pathology VQA via Token Compression

2025-07-19 · Weimin Lyu, Qingqiao Hu, Kehan Qi, Zhan Shi 외 arxiv

Whole-slide images (WSIs) in pathology can reach up to 10,000 x 10,000 pixels, posing significant challenges for multimodal large language model (MLLM) due to long context length and high computational demands. Previous …

Visual Question AnsweringAnswer Generation

SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding

2024-10-15 · CVPR 2025 1 · Ying Chen, Guoan Wang, Yuanfeng Ji, Yanjun Li 외

Despite the progress made by multimodal large language models (MLLMs) in computational pathology, they remain limited by a predominant focus on patch-level analysis, missing essential contextual information at the whole-…

Instruction FollowingVisual Question Answering (VQA)whole slide images

Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images

2026-02-23 · Yuxuan Yang, Zhonghao Yan, Yi Zhang, Bo Yun 외 arxiv

Hepatocellular Carcinoma diagnosis relies heavily on the interpretation of gigapixel Whole Slide Images. However, current computational approaches are constrained by fixed-resolution processing mechanisms and inefficient…

A Multimodal Knowledge-enhanced Whole-slide Pathology Foundation Model

2024-07-22 · Yingxue Xu, Yihui Wang, Fengtao Zhou, Jiabo Ma 외

Remarkable strides in computational pathology have been made in the task-agnostic foundation model that advances the performance of a wide array of downstream clinical tasks. Despite the promising performance, there are …

Diagnosticwhole slide images

MPath: Multimodal Pathology Report Generation from Whole Slide Images

2025-12-10 · Noorul Wahab, Nasir Rajpoot arxiv

Automated generation of diagnostic pathology reports directly from whole slide images (WSIs) is an emerging direction in computational pathology. Translating high-resolution tissue patterns into clinically coherent text …