paper-with-me

홈 › Papers

A benchmark multimodal oro-dental dataset for large vision-language models

2025-11-07 · Haoxin Lv, Ijazul Haq, Jin Du, Jiaxin Ma, Binnian Zhu, Xiaobing Dang, Chaoan Liang, Ruxu Du, Yingjie Zhang, Muhammad Saqib arxiv

The advancement of artificial intelligence in oral healthcare relies on the availability of large-scale multimodal datasets that capture the complexity of clinical practice. In this paper, we present a comprehensive multimodal dataset, comprising 8775 dental checkups from 4800 patients collected over eight years (2018-2025), with patients ranging from 10 to 90 years of age. The dataset includes 50000 intraoral images, 8056 radiographs, and detailed textual records, including diagnoses, treatment plans, and follow-up notes. The data were collected under standard ethical guidelines and annotated for benchmarking. To demonstrate its utility, we fine-tuned state-of-the-art large vision-language models, Qwen-VL 3B and 7B, and evaluated them on two tasks: classification of six oro-dental anomalies and generation of complete diagnostic reports from multimodal inputs. We compared the fine-tuned models with their base counterparts and GPT-4o. The fine-tuned models achieved substantial gains over these baselines, validating the dataset and underscoring its effectiveness in advancing AI-driven oro-dental healthcare solutions. The dataset is publicly available, providing an essential resource for future research in AI dentistry.

📄 PDF Abstract BibTeX arXiv:2511.04948

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OralGPT-Omni: A Versatile Dental Multimodal Large Language Model

2025-11-27 · Jing Hao, Yuci Liang, Lizhuo Lin, Yuxuan Fan 외 arxiv

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annot…

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

2025-12-12 · Zhenyang Cai, Jiaming Zhang, Junjie Zhao, Ziyi Zeng 외 arxiv

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack suf…

Reinforcement Learning

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

2026-05-28 · Kai Bian, Xucheng Guo, Bin Chen, Lingyan Ruan 외 arxiv

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outsid…

Question Answering

Large AI Models in Dental Healthcare: From General-Purpose Systems to Domain-Specific Foundation Models

2026-06-01 · Sema Helali, Lina Abu Nada, Sausan Al Kawas, Alaa Abd-Alrazaq 외 arxiv

Background: Oral diseases affect nearly 3.5 billion people worldwide, yet the comparative clinical potential of large-scale AI models in dentistry remains poorly understood. Three distinct model categories have emerged: …

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice

2026-05-02 · Rongyang Wang, Shuang Zhou, Jiashuo Wang, Wenya Xie 외 arxiv

Multimodal large language models (MLLMs) have emerged as a promising paradigm for dental image analysis. However, their ability to capture the multi-level cognitive processes required for radiographic analysis remains un…