paper-with-me

Papers

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

2026-02-13 · Baorong Shi, Bo Cui, Boyuan Jiang, Deli Yu, Fang Qian, Haihua Yang, Huichao Wang, Jiale Chen, Jianfei Pan, Jieqiong Cao, Jinghao Lin, Kai Wu, Lin Yang, Shengsheng Yao, Tao Chen, Xiaojun Xiao, Xiaozhong Ji, Xu Wang, Yijun He, Zhixiong Yang arxiv

We present MedXIAOHE, a medical vision-language foundation model designed to advance general-purpose medical understanding and reasoning in real-world clinical applications. MedXIAOHE achieves state-of-the-art performance across diverse medical benchmarks and surpasses leading closed-source multimodal systems on multiple capabilities. To achieve this, we propose an entity-aware continual pretraining framework that organizes heterogeneous medical corpora to broaden knowledge coverage and reduce long-tail gaps (e.g., rare diseases). For medical expert-level reasoning and interaction, MedXIAOHE incorporates diverse medical reasoning patterns via reinforcement learning and tool-augmented agentic training, enabling multi-step diagnostic reasoning with verifiable decision traces. To improve reliability in real-world use, MedXIAOHE integrates user-preference rubrics, evidence-grounded reasoning, and low-hallucination long-form report generation, with improved adherence to medical instructions. We release this report to document our practical design choices, scaling insights, and evaluation framework, hoping to inspire further research.

📄 PDF Abstract BibTeX arXiv:2602.12705

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningContinual Pretraining

Similar Papers 제목 키워드 기반

Multimodal Large Language Models for Medicine: A Comprehensive Survey

2025-04-29 · Jiarui Ye, Hao Tang

MLLMs have recently become a focal point in the field of artificial intelligence research. Building on the strong capabilities of LLMs, MLLMs are adept at addressing complex multi-modal tasks. With the release of GPT-4, …

Medical DiagnosisSurvey

ChEF: A Comprehensive Evaluation Framework for Standardized Assessment of Multimodal Large Language Models

2023-11-05 · Zhelun Shi, Zhipin Wang, Hongxing Fan, Zhenfei Yin 외

Multimodal Large Language Models (MLLMs) have shown impressive abilities in interacting with visual content with myriad potential downstream tasks. However, even though a list of benchmarks has been proposed, the capabil…

HallucinationIn-Context LearningInstruction FollowingQuestion Answering

FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation

2024-09-27 · Yuki Imajuku, Yoko Yamakata, Kiyoharu Aizawa

Research on food image understanding using recipe data has been a long-standing focus due to the diversity and complexity of the data. Moreover, food is inextricably linked to people's lives, making it a vital research a…

Recipe Generation

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

2025-08-22 · Ruiqi Wu, Yuang Yao, Tengfei Ma, Chenran Zhang 외 arxiv

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical doma…

Reinforcement LearningMultimodal Reasoning

A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine

2024-05-14 · Hanguang Xiao, Feizhong Zhou, Xingyue Liu, Tianqi Liu 외

Since the release of ChatGPT and GPT-4, large language models (LLMs) and multimodal large language models (MLLMs) have attracted widespread attention for their exceptional capabilities in understanding, reasoning, and ge…

Survey