paper-with-me

홈 › Papers

Generating High Quality Synthetic Data for Dutch Medical Conversations

2026-03-25 · Cecilia Kuan, Aditya Kamlesh Parikh, Henk van den Heuvel arxiv

Medical conversations offer insights into clinical communication often absent from Electronic Health Records. However, developing reliable clinical Natural Language Processing (NLP) models is hampered by the scarcity of domain-specific datasets, as clinical data are typically inaccessible due to privacy and ethical constraints. To address these challenges, we present a pipeline for generating synthetic Dutch medical dialogues using a Dutch fine-tuned Large Language Model, with real medical conversations serving as linguistic and structural reference. The generated dialogues were evaluated through quantitative metrics and qualitative review by native speakers and medical practitioners. Quantitative analysis revealed strong lexical variety and overly regular turn-taking, suggesting scripted rather than natural conversation flow. Qualitative review produced slightly below-average scores, with raters noting issues in domain specificity and natural expression. The limited correlation between quantitative and qualitative results highlights that numerical metrics alone cannot fully capture linguistic quality. Our findings demonstrate that generating synthetic Dutch medical dialogues is feasible but requires domain knowledge and carefully structured prompting to balance naturalness and structure in conversation. This work provides a foundation for expanding Dutch clinical NLP resources through ethically generated synthetic data.

📄 PDF Abstract BibTeX arXiv:2604.09645

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GEITje 7B Ultra: A Conversational Model for Dutch

2024-12-05 · Bram Vanroy

Language models have rapidly evolved, predominantly focusing on English while often neglecting extensive pretraining in other languages. This approach has required initiatives to adapt powerful, English-centric models to…

model

Towards Fairness Assessment of Dutch Hate Speech Detection

2025-06-14 · Julie Bauer, Rishabh Kaushal, Thales Bertaglia, Adriana Iamnitchi

Numerous studies have proposed computational methods to detect hate speech online, yet most focus on the English language and emphasize model development. In this study, we evaluate the counterfactual fairness of hate sp…

counterfactualFairnessHate Speech DetectionSentence

Language Resources for Dutch Large Language Modelling

2023-12-20 · Bram Vanroy

Despite the rapid expansion of types of large language models, there remains a notable gap in models specifically designed for the Dutch language. This gap is not only a shortage in terms of pretrained Dutch models but a…

Language Modelling

MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch

2025-09-15 · Nikolay Banar, Ehsan Lotfi, Jens Van Nooten, Cristina Arhiliuc 외 arxiv

Recently, embedding resources, including models, benchmarks, and datasets, have been widely released to support a variety of languages. However, the Dutch language remains underrepresented, typically comprising only a sm…

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

2026-04-01 · Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy 외 arxiv

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present i…