paper-with-me

홈 › Papers

Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning

2025-05-21 · Zhe Xu, Cheng Jin, Yihui Wang, Ziyi Liu, Hao Chen

Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, existing methods exhibit limited reasoning capabilities, which hamper their ability to handle complex diagnostic scenarios. Additionally, the enormous size of pathological images leads to severe computational burdens, further restricting their practical deployment. To address these limitations, we introduce a novel bilateral reinforcement learning framework comprising two synergistic branches. One reinforcement branch enhances the reasoning capability by enabling the model to learn task-specific decision processes, i.e., pathology rationales, directly from labels without explicit reasoning supervision. While the other branch dynamically allocates a tailored number of tokens to different images based on both their visual content and task context, thereby optimizing computational efficiency. We apply our method to various pathological tasks such as visual question answering, cancer subtyping, and lesion detection. Extensive experiments show an average +41.7 absolute performance improvement with 70.3% lower inference costs over the base models, achieving both reasoning accuracy and computational efficiency.

📄 PDF Abstract BibTeX arXiv:2505.15687

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyDiagnosticLesion DetectionQuestion AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

2026-06-14 · Zebang Cheng, Shuimu Chen, Boxue Yang, Yuanshen Guan 외 arxiv

Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounced in human-centered scenarios involving …

Reinforcement Learning

Reinforced Attention Learning

2026-02-04 · Bangzheng Li, Jianmo Ni, Chen Qu, Ian Miao 외 arxiv

Post-training with Reinforcement Learning (RL) has substantially improved reasoning in Large Language Models (LLMs) via test-time scaling. However, extending this paradigm to Multimodal LLMs (MLLMs) through verbose ratio…

Reinforcement LearningKnowledge Distillation

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

2026-06-08 · Yutong Bian, Dongjie Cheng, Heming Xia, Yongqi Li 외 arxiv

Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves from text-based multimodal reasoning towar…

Multimodal Reasoning

AGIR: Assessing 3D Gait Impairment with Reasoning based on LLMs

2025-03-23 · Diwei Wang, Cédric Bobenrieth, Hyewon Seo

Assessing gait impairment plays an important role in early diagnosis, disease monitoring, and treatment evaluation for neurodegenerative diseases. Despite its widespread use in clinical practice, it is limited by subject…

Large Language Model

Multimodal Model for Computational Pathology:Representation Learning and Image Compression

2026-03-19 · Peihang Wu, Zehong Chen, Lijian Xu arxiv

Whole slide imaging (WSI) has transformed digital pathology by enabling computational analysis of gigapixel histopathology images. Recent foundation model advances have accelerated progress in computational pathology, fa…

Representation LearningImage CompressionFew-Shot Learning