paper-with-me

Papers

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

2025-11-02 · Yan Shu, Chi Liu, Robin Chen, Derek Li, Bryan Dai arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly focused on empowering MLLMs with medical conversational abilities, which hold significant promise for clinical applications. However, medical data presents unique challenges due to its heterogeneous nature -- encompassing diverse modalities including 2D images, 3D volumetric scans, and temporal video sequences. The substantial domain gap and data format inconsistencies across these modalities have hindered the development of unified medical MLLMs. To address these challenges, we propose Fleming-VL, a unified end-to-end framework for comprehensive medical visual understanding across heterogeneous modalities. Fleming-VL tackles this problem from a data-centric perspective through three key strategies: (1) scaling up pretraining by integrating long-context data from both natural and medical-specific domains; (2) complementing fine-tuning with rare medical data, including holistic video analysis and underrepresented 2D modalities such as ultrasound and dermoscopy images; (3) extending existing evaluation frameworks to incorporate 3D volumetric and video understanding benchmarks. Through supervised fine-tuning (SFT) and group relative policy optimization (GRPO), we develop Fleming-VL in multiple model scales. Extensive experiments demonstrate that Fleming-VL achieves state-of-the-art performance across multiple benchmarks, including medical VQA, video QA, and 3D medical image understanding. We publicly release Fleming-VL to promote transparent, reproducible, and auditable progress in medical AI.

📄 PDF Abstract BibTeX arXiv:2511.00916

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringImage CaptioningVisual Reasoning

Similar Papers 제목 키워드 기반

MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

2026-02-06 · Ankan Deria, Komal Kumar, Adinath Madhavrao Dukre, Eran Segal 외 arxiv

Multimodal large language models have advanced rapidly, but their adoption in medicine is constrained by limited domain coverage, imperfect modality alignment, and insufficient grounded reasoning. We introduce MedMO, a m…

Medical Report GenerationReinforcement Learning

Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning

2025-09-18 · Chi Liu, Derek Li, Yan Shu, Robin Chen 외 arxiv

While large language models show promise in medical applications, achieving expert-level clinical reasoning remains challenging due to the need for both accurate answers and transparent reasoning processes. To address th…

Reinforcement Learning

V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval

2026-02-05 · Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao 외 arxiv

Multimodal Large Language Models (MLLMs) have recently been applied to universal multimodal retrieval, where Chain-of-Thought (CoT) reasoning improves candidate reranking. However, existing approaches remain largely lang…

Reinforcement Learning

Generative Universal Verifier as Multimodal Meta-Reasoner

2025-10-15 · Xinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao 외 arxiv

We introduce Generative Universal Verifier, a novel concept and plugin designed for next-generation multimodal reasoning in vision-language models and unified multimodal models, providing the fundamental capability of re…

Multimodal ReasoningImage Generation

Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning

2026-06-30 · Kaitao Chen, Weiqian Zhao, Jiamin Wu, Qihao Zheng 외 arxiv

Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle with medical images, which typically exhibit extremely sparse visual evidence to …

Reinforcement LearningMultimodal ReasoningQuestion Answering