paper-with-me

홈 › Papers

ReinPath: A Multimodal Reinforcement Learning Approach for Pathology

2026-01-21 · Kangcheng Zhou, Jun Jiang, Qing Zhang, Shuang Zheng, Qingli Li, Shugong Xu arxiv

Interpretability is significant in computational pathology, leading to the development of multimodal information integration from histopathological image and corresponding text data.However, existing multimodal methods have limited interpretability due to the lack of high-quality dataset that support explicit reasoning and inference and simple reasoning process.To address the above problems, we introduce a novel multimodal pathology large language model with strong reasoning capabilities.To improve the generation of accurate and contextually relevant textual descriptions, we design a semantic reward strategy integrated with group relative policy optimization.We construct a high-quality pathology visual question answering (VQA) dataset, specifically designed to support complex reasoning tasks.Comprehensive experiments conducted on this dataset demonstrate that our method outperforms state-of-the-art methods, even when trained with only 20% of the data.Our method also achieves comparable performance on downstream zero-shot image classification task compared with CLIP.

📄 PDF Abstract BibTeX arXiv:2601.14757

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-Shot Image ClassificationVisual Question AnsweringReinforcement Learning

Similar Papers 제목 키워드 기반

Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner

2025-05-16 · Wenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng 외

Recent advances in vision language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging subdomain, with current pathology specific VLMs exhibiting li…

Cross-Modal RetrievalDiagnosticImage DescriptionMultimodal Reasoning+5

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

2025-08-04 · Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang 외 arxiv

Although Vision Language Models (VLMs) have shown strong generalization in medical imaging, pathology presents unique challenges due to ultra-high resolution, complex tissue structures, and nuanced clinical semantics. Th…

Visual Question AnsweringReinforcement Learning

Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning

2025-05-21 · Zhe Xu, Cheng Jin, Yihui Wang, Ziyi Liu 외

Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, exis…

Computational EfficiencyDiagnosticLesion DetectionQuestion Answering+1

TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots

2025-11-20 · Tianyu Liu, Weihao Xuan, Hao Wu, Peter Humphrey 외 arxiv

Advances in AI have introduced several strong models in computational pathology to usher it into the era of multi-modal diagnosis, analysis, and interpretation. However, the current pathology-specific visual language mod…

Reinforcement Learning

DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis

2025-07-24 · Minxi Ouyang, Lianghui Zhu, Yaqing Bao, Qiang Huang 외 arxiv

Multimodal large models have shown great potential in automating pathology image analysis. However, current multimodal models for gastrointestinal pathology are constrained by both data quality and reasoning transparency…

Reinforcement Learning