paper-with-me

Papers

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

2025-12-12 · Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng, Yanchao Li, Jingyi Liang, Junying Chen, Yunjin Yang, Jiajun You, Shuzhi Deng, Tongfei Wang, Wanting Chen, Chunxiu Hao, Ruiqi Xie, Zhenwei Wen, Xiangyi Feng, Zou Ting, Jin Zou Lin, Jianquan Li, Guangjun Yu, Liangyi Chen, Junwen Wang, Shan Jiang, Benyou Wang arxiv

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack sufficient reasoning ability for precise diagnosis. To address these limitations, we present DentalGPT, a specialized dental MLLM developed through high-quality domain knowledge injection and reinforcement learning. Specifically, the largest annotated multimodal dataset for dentistry to date was constructed by aggregating over 120k dental images paired with detailed descriptions that highlight diagnostically relevant visual features, making it the multimodal dataset with the most extensive collection of dental images to date. Training on this dataset significantly enhances the MLLM's visual understanding of dental conditions, while the subsequent reinforcement learning stage further strengthens its capability for multimodal complex reasoning. Comprehensive evaluations on intraoral and panoramic benchmarks, along with dental subsets of medical VQA benchmarks, show that DentalGPT achieves superior performance in disease classification and dental VQA tasks, outperforming many state-of-the-art MLLMs despite having only 7B parameters. These results demonstrate that high-quality dental data combined with staged adaptation provides an effective pathway for building capable and domain-specialized dental MLLMs.

📄 PDF Abstract BibTeX arXiv:2512.11558

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

2025-11-27 · Jing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan 외 arxiv

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annot…

ChatGPT for Shaping the Future of Dentistry: The Potential of Multi-Modal Large Language Model

2023-03-23 · Hanyao Huang, Ou Zheng, Dongdong Wang, Jiayi Yin 외

The ChatGPT, a lite and conversational variant of Generative Pretrained Transformer 4 (GPT-4) developed by OpenAI, is one of the milestone Large Language Models (LLMs) with billions of parameters. LLMs have stirred up mu…

Language ModelingLanguage ModellingLarge Language Model

Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

2025-03-09 · Wenxuan Huang, Bohan Jia, Zijie Zhai, Shaosheng Cao 외

DeepSeek-R1-Zero has successfully demonstrated the emergence of reasoning capabilities in LLMs purely through Reinforcement Learning (RL). Inspired by this breakthrough, we explore how RL can be utilized to enhance the r…

MathMultimodal ReasoningReinforcement Learning (RL)

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

2026-01-29 · Wenxuan Huang, Yu Zeng, Qiuchen Wang, Zhen Fang 외 arxiv

Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting M…

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

2026-08-03 · Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan 외 hf

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world pro…

Reinforcement Learning