paper-with-me

홈 › Papers

Scaling Arabic Medical Chatbots Using Synthetic Data: Enhancing Generative AI with Synthetic Patient Records

2025-09-12 · Abdulrahman Allam, Seif Ahmed, Ali Hamdi, Khaled Shaban arxiv

The development of medical chatbots in Arabic is significantly constrained by the scarcity of large-scale, high-quality annotated datasets. While prior efforts compiled a dataset of 20,000 Arabic patient-doctor interactions from social media to fine-tune large language models (LLMs), model scalability and generalization remained limited. In this study, we propose a scalable synthetic data augmentation strategy to expand the training corpus to 100,000 records. Using advanced generative AI systems ChatGPT-4o and Gemini 2.5 Pro we generated 80,000 contextually relevant and medically coherent synthetic question-answer pairs grounded in the structure of the original dataset. These synthetic samples were semantically filtered, manually validated, and integrated into the training pipeline. We fine-tuned five LLMs, including Mistral-7B and AraGPT2, and evaluated their performance using BERTScore metrics and expert-driven qualitative assessments. To further analyze the effectiveness of synthetic sources, we conducted an ablation study comparing ChatGPT-4o and Gemini-generated data independently. The results showed that ChatGPT-4o data consistently led to higher F1-scores and fewer hallucinations across all models. Overall, our findings demonstrate the viability of synthetic augmentation as a practical solution for enhancing domain-specific language models in-low resource medical NLP, paving the way for more inclusive, scalable, and accurate Arabic healthcare chatbot systems.

📄 PDF Abstract BibTeX arXiv:2509.10108

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Arabic Chatbot Technologies in Education: An Overview

2025-09-04 · Hicham Bourhil, Yacine El Younoussi arxiv

The recent advancements in Artificial Intelligence (AI) in general, and in Natural Language Processing (NLP) in particular, and some of its applications such as chatbots, have led to their implementation in different dom…

Improving Domain Independent Question Parsing with Synthetic Treebanks

2018-08-01 · COLING 2018 8 · Halim-Antoine Boukaram, Nizar Habash, Micheline Ziadee, Majd Sakr

Automatic syntactic parsing for question constructions is a challenging task due to the paucity of training examples in most treebanks. The near absence of question constructions is due to the dominance of the news domai…

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

2026-02-17 · Ahmed Khaled Khamis, Hesham Ali arxiv

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …

Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to Speech

Taxonomy of AISecOps Threat Modeling for Cloud Based Medical Chatbots

2023-05-18 · Ruby Annette J, Aisha Banu, Sharon Priya S, Subash Chandran

Artificial Intelligence (AI) is playing a vital role in all aspects of technology including cyber security. Application of Conversational AI like the chatbots are also becoming very popular in the medical field to provid…

Chatbot

New Arabic Medical Dataset for Diseases Classification

2021-06-29 · Jaafar Hammoud, Aleksandra Vatian, Natalia Dobrenko, Nikolai Vedernikov 외

The Arabic language suffers from a great shortage of datasets suitable for training deep learning models, and the existing ones include general non-specialized classifications. In this work, we introduce a new Arab medic…

Classification