paper-with-me

홈 › Papers

MedINST: Meta Dataset of Biomedical Instructions

2024-10-17 · Wenhan Han, Meng Fang, Zihan Zhang, Yu Yin, Zirui Song, Ling Chen, Mykola Pechenizkiy, Qingyu Chen

The integration of large language model (LLM) techniques in the field of medical analysis has brought about significant advancements, yet the scarcity of large, diverse, and well-annotated datasets remains a major challenge. Medical data and tasks, which vary in format, size, and other parameters, require extensive preprocessing and standardization for effective use in training LLMs. To address these challenges, we introduce MedINST, the Meta Dataset of Biomedical Instructions, a novel multi-domain, multi-task instructional meta-dataset. MedINST comprises 133 biomedical NLP tasks and over 7 million training samples, making it the most comprehensive biomedical instruction dataset to date. Using MedINST as the meta dataset, we curate MedINST32, a challenging benchmark with different task difficulties aiming to evaluate LLMs' generalization ability. We fine-tune several LLMs on MedINST and evaluate on MedINST32, showcasing enhanced cross-task generalization.

📄 PDF Abstract BibTeX arXiv:2410.13458

Code (1)

aialt/medinst 공식 구현

Tasks

Language ModelingLanguage ModellingLarge Language Model

Similar Papers 제목 키워드 기반

In-BoXBART: Get Instructions into Biomedical Multi-task Learning

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Single-task models have proven pivotal in solving specific tasks; however, they have limitations in real-world applications where multi-tasking is necessary and domain shifts are exhibited. Recently, instructional prompt…

Few-Shot LearningMulti-Task Learning

In-BoXBART: Get Instructions into Biomedical Multi-Task Learning

2022-04-15 · Findings (NAACL) 2022 7 · Mihir Parmar, Swaroop Mishra, Mirali Purohit, Man Luo 외

Single-task models have proven pivotal in solving specific tasks; however, they have limitations in real-world applications where multi-tasking is necessary and domain shifts are exhibited. Recently, instructional prompt…

Few-Shot LearningMulti-Task Learning

AlpaCare:Instruction-tuned Large Language Models for Medical Application

2023-10-23 · Xinlu Zhang, Chenxin Tian, Xianjun Yang, Lichang Chen 외

Instruction-finetuning (IFT) has become crucial in aligning Large Language Models (LLMs) with diverse human needs and has shown great potential in medical applications. However, previous studies mainly fine-tune LLMs on …

DiversityInstruction Following

Labeling instructions matter in biomedical image analysis

2022-07-20 · Tim Rädsch, Annika Reinke, Vivienn Weru, Minu D. Tizabi 외

Biomedical image analysis algorithm validation depends on high-quality annotation of reference datasets, for which labeling instructions are key. Despite their importance, their optimization remains largely unexplored. H…

BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing

2022-06-30 · Jason Alan Fries, Leon Weber, Natasha Seelam, Gabriel Altay 외

Training and evaluating language models increasingly requires the construction of meta-datasets --diverse collections of curated data with clear provenance. Natural language prompting has recently lead to improved zero-s…

DiversityLanguage Model EvaluationLanguage ModelingLanguage Modelling+4