paper-with-me

Papers

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

2025-11-28 · Timothy Ossowski, Sheng Zhang, Qianchu Liu, Guanghui Qin, Reuben Tan, Tristan Naumann, Junjie Hu, Hoifung Poon arxiv

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tasks. We investigate strategies for training and data curation to develop a robust multimodal reasoning model in the medical domain. Our work focuses on supervised fine-tuning (SFT) and explores data recipes that leverage structured reasoning traces. Using our proposed data recipe, we scale experiments to a dataset of over 8 million examples and 6.8 billion response tokens, achieving state-of-the-art performance among open-source models across diverse out-of-distribution medical benchmark tasks. Our results further indicate that curating a high-quality, diverse training dataset with varying structured reasoning trace lengths enables the fine-tuned model to self-calibrate its reasoning trajectory lengths based on the downstream task, without explicit supervision. We present key insights, describe the data curation strategy, and outline next steps toward developing robust medical vision-language reasoning system.

📄 PDF Abstract BibTeX arXiv:2511.23269

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains

2026-03-01 · Ahmadreza Jeddi, Kimia Shaban, Negin Baghbanzadeh, Natasha Sharan 외 arxiv

Reinforcement learning (RL) is increasingly used to post-train medical Vision-Language Models (VLMs), yet it remains unclear whether RL improves medical visual reasoning or mainly sharpens behaviors already induced by su…

Reinforcement LearningVisual Reasoning

MedVLThinker: Simple Baselines for Multimodal Medical Reasoning

2025-08-04 · Xiaoke Huang, Juncheng Wu, Hui Liu, Xianfeng Tang 외 arxiv

Large Reasoning Models (LRMs) have introduced a new paradigm in AI by enabling models to ``think before responding" via chain-of-thought reasoning. However, the absence of open and reproducible recipes for building reaso…

Reinforcement Learning

RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes

2018-09-04 · EMNLP 2018 10 · Semih Yagcioglu, Aykut Erdem, Erkut Erdem, Nazli Ikizler-Cinbis

Understanding and reasoning about cooking recipes is a fruitful research direction towards enabling machines to interpret procedural text. In this work, we introduce RecipeQA, a dataset for multimodal comprehension of co…

Reading Comprehension

Aloe-Vision: Robust Vision-Language Models for Healthcare

2026-06-25 · Jaume Guasch-Martí, Enrique Lopez-Cuena, Martín Suárez-Fernández, Jordi Bayarri-Planas 외 arxiv

Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the…

A Recipe for Creating Multimodal Aligned Datasets for Sequential Tasks

2020-05-19 · ACL 2020 6 · Angela S. Lin, Sudha Rao, Asli Celikyilmaz, Elnaz Nouri 외

Many high-level procedural tasks can be decomposed into sequences of instructions that vary in their order and choice of tools. In the cooking domain, the web offers many partially-overlapping text and video recipes (i.e…

Descriptive