paper-with-me

홈 › Papers

PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis

2026-01-12 · Jiao Xu, Junwei Liu, Jiangwei Lao, Qi Zhu, Yunpeng Zhao, Congyun Jin, Shinan Liu, Zhihong Lu, Lihe Zhang, Xin Chen, Jian Wang, Ping Wang arxiv

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous inputs and require ongoing contextual understanding during patient-physician interactions. To bridge this gap, we introduce PulseMind, a new family of multi-modal diagnostic models that integrates a systematically curated dataset, a comprehensive evaluation benchmark, and a tailored training framework. Specifically, we first construct a diagnostic dataset, MediScope, which comprises 98,000 real-world multi-turn consultations and 601,500 medical images, spanning over 10 major clinical departments and more than 200 sub-specialties. Then, to better reflect the requirements of real-world clinical diagnosis, we develop the PulseMind Benchmark, a multi-turn diagnostic consultation benchmark with a four-dimensional evaluation protocol comprising proactiveness, accuracy, usefulness, and language quality. Finally, we design a training framework tailored for multi-modal clinical diagnostics, centered around a core component named Comparison-based Reinforcement Policy Optimization (CRPO). Compared to absolute score rewards, CRPO uses relative preference signals from multi-dimensional com-parisons to provide stable and human-aligned training guidance. Extensive experiments demonstrate that PulseMind achieves competitive performance on both the diagnostic consultation benchmark and public medical benchmarks.

📄 PDF Abstract BibTeX arXiv:2601.07344

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

2026-07-10 · Runhan Shi, Quan Zhou, Yuqian Xu, Shuai Yang 외 arxiv

Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators…

Response Generation

Weakly-Supervised Multimodal Learning on MIMIC-CXR

2024-11-15 · Andrea Agostini, Daphné Chopard, Yang Meng, Norbert Fortin 외

Multimodal data integration and label scarcity pose significant challenges for machine learning in medical settings. To address these issues, we conduct an in-depth evaluation of the newly proposed Multimodal Variational…

Data IntegrationMixture-of-Experts

A Demonstration of Adaptive Collaboration of Large Language Models for Medical Decision-Making

2024-10-31 · Yubin Kim, Chanwoo Park, Hyewon Jeong, Cristina Grau-Vilchez 외

Medical Decision-Making (MDM) is a multi-faceted process that requires clinicians to assess complex multi-modal patient data patient, often collaboratively. Large Language Models (LLMs) promise to streamline this process…

Decision MakingDiagnostic

A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning

2025-11-17 · Liuyi Jin, Pasan Gunawardena, Amran Haroon, Runzhi Wang 외 arxiv

Emergency Medical Technicians (EMTs) operate in high-pressure environments, making rapid, life-critical decisions under heavy cognitive and operational loads. We present EMSGlass, a smart-glasses system powered by EMSNet…

Med-MMFL: A Multimodal Federated Learning Benchmark in Healthcare

2026-02-04 · Aavash Chhetri, Bibek Niroula, Pratik Shrestha, Yash Raj Shrestha 외 arxiv

Federated learning (FL) enables collaborative model training across decentralized medical institutions while preserving data privacy. However, medical FL benchmarks remain scarce, with existing efforts focusing mainly on…

Federated Learning