paper-with-me

홈 › Papers

MedInjection-FR: Exploring the Role of Native, Synthetic, and Translated Data in Biomedical Instruction Tuning

2026-03-06 · Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils, Benoit Favre, Richard Dufour arxiv

Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts. Yet, in specialized fields such as medicine, the scarcity of high-quality French instruction data limits effective supervision. To address this gap, we introduce MedInjection-FR, a large-scale French biomedical instruction dataset comprising 571K instruction-response pairs drawn from three complementary sources: native, synthetic, and translated data. We design a controlled experimental framework to systematically assess how data provenance affects instruction tuning, using Qwen-4B-Instruct fine-tuned across seven configurations combining these sources. Results show that native data yield the strongest performance, while mixed setups, particularly native and translated, provide complementary benefits. Synthetic data alone remains less effective but contributes positively when balanced with native supervision. Evaluation on open-ended QA combines automatic metrics, LLM-as-a-judge assessment, and human expert review; although LLM-based judgments correlate best with human ratings, they show sensitivity to verbosity. These findings highlight that data authenticity and diversity jointly shape downstream adaptation and that heterogeneous supervision can mitigate the scarcity of native French medical instructions.

📄 PDF Abstract BibTeX arXiv:2603.06905

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Tagged Back-Translation

2019-06-15 · WS 2019 8 · Isaac Caswell, Ciprian Chelba, David Grangier

Recent work in Neural Machine Translation (NMT) has shown significant quality gains from noised-beam decoding during back-translation, a method to generate synthetic parallel data. We show that the main role of such synt…

Machine TranslationNMTTranslation

From Synthetic to Native: Benchmarking Multilingual Intent Classification in Logistics Customer Service

2026-03-24 · Haoyu He, Jinyu Zhuang, Haoran Chu, Shuhang Yu 외 arxiv

Multilingual intent classification is central to customer-service systems on global logistics platforms, where models must process noisy user queries across languages and hierarchical label spaces. Yet most existing mult…

Cross-Lingual TransferIntent Classification

Exploring the Role of Synthetic Data Augmentation in Controllable Human-Centric Video Generation

2026-04-23 · Yuanchen Fei, Yude Zou, Zejian Kang, Ming Li 외 arxiv

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity o…

Data AugmentationVideo Generation

Exploring Gap Filling as a Cheaper Alternative to Reading Comprehension Questionnaires when Evaluating Machine Translation for Gisting

2018-09-02 · WS 2018 10 · Mikel L. Forcada, Carolina Scarton, Lucia Specia, Barry Haddow 외

A popular application of machine translation (MT) is gisting: MT is consumed as is to make sense of text in a foreign language. Evaluation of the usefulness of MT for gisting is surprisingly uncommon. The classical metho…

Machine TranslationReading ComprehensionSentenceTranslation

A Corpus of Native, Non-native and Translated Texts

2016-05-01 · LREC 2016 5 · Sergiu Nisioi, Ella Rabinovich, Liviu P. Dinu, Shuly Wintner

We describe a monolingual English corpus of original and (human) translated texts, with an accurate annotation of speaker properties, including the original language of the utterances and the speaker{'}s country of origi…