paper-with-me

홈 › Papers

ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling

2025-07-21 · Yuanhe Tian, Junjie Liu, Zhizhou Kou, Yuxiang Li, Yan Song arxiv

Building high-quality data resources is crucial for advancing artificial intelligence research and applications in specific domains, particularly in the Chinese medical domain. Existing Chinese medical datasets are limited in size and narrow in domain coverage, falling short of the diverse corpora required for effective pre-training. Moreover, most datasets are designed solely for LLM fine-tuning and do not support pre-training and reinforcement learning from human feedback (RLHF). In this paper, we propose a Chinese medical dataset named ChiMed 2.0, which extends our previous work ChiMed, and covers data collected from Chinese medical online platforms and generated by LLMs. ChiMed 2.0 contains 204.4M Chinese characters covering both traditional Chinese medicine classics and modern general medical data, where there are 164.8K documents for pre-training, 351.6K question-answering pairs for supervised fine-tuning (SFT), and 41.7K preference data tuples for RLHF. To validate the effectiveness of our approach for training a Chinese medical LLM, we conduct further pre-training, SFT, and RLHF experiments on representative general domain LLMs and evaluate their performance on medical benchmark datasets. The results show performance gains across different model scales, validating the dataset's effectiveness and applicability.

📄 PDF Abstract BibTeX arXiv:2507.15275

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

ChiMed-GPT: A Chinese Medical Large Language Model with Full Training Regime and Better Alignment to Human Preferences

2023-11-10 · Yuanhe Tian, Ruyi Gan, Yan Song, Jiaxing Zhang 외

Recently, the increasing demand for superior medical services has highlighted the discrepancies in the medical infrastructure. With big data, especially texts, forming the foundation of medical services, there is an exig…

Dialogue GenerationLanguage ModelingLanguage ModellingLarge Language Model+1

ChiMed: A Chinese Medical Corpus for Question Answering

2019-08-01 · WS 2019 8 · Yuanhe Tian, Weicheng Ma, Fei Xia, Yan Song

Question answering (QA) is a challenging task in natural language processing (NLP), especially when it is applied to specific domains. While models trained in the general domain can be adapted to a new target domain, the…

Question Answering

Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare

2023-10-27 · Junling Liu, ZiMing Wang, Qichen Ye, Dading Chong 외

Large Language Models (LLMs) have introduced a new era of proficiency in comprehending complex healthcare and biomedical topics. However, there is a noticeable lack of models in languages other than English and models th…

Language ModelingLanguage Modelling

Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks

2025-10-13 · Wenya Xie, Qingying Xiao, Yu Zheng, Xidong Wang 외 arxiv

The rise of large language models (LLMs) has transformed healthcare by offering clinical guidance, yet their direct deployment to patients poses safety risks due to limited domain expertise. To mitigate this, we propose …

Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model

2023-10-13 · Qichen Ye, Junling Liu, Dading Chong, Peilin Zhou 외

Integrating large language models (LLMs) into healthcare holds great potential but faces challenges. Pre-training LLMs from scratch for domains like medicine is resource-heavy and often unfeasible. On the other hand, sol…

Knowledge GraphsLanguage ModelingLanguage ModellingLarge Language Model+4