paper-with-me

홈 › Papers

MING-MOE: Enhancing Medical Multi-Task Learning in Large Language Models with Sparse Mixture of Low-Rank Adapter Experts

2024-04-13 · Yusheng Liao, Shuyang Jiang, Yu Wang, Yanfeng Wang

Large language models like ChatGPT have shown substantial progress in natural language understanding and generation, proving valuable across various disciplines, including the medical field. Despite advancements, challenges persist due to the complexity and diversity inherent in medical tasks which often require multi-task learning capabilities. Previous approaches, although beneficial, fall short in real-world applications because they necessitate task-specific annotations at inference time, limiting broader generalization. This paper introduces MING-MOE, a novel Mixture-of-Expert~(MOE)-based medical large language model designed to manage diverse and complex medical tasks without requiring task-specific annotations, thus enhancing its usability across extensive datasets. MING-MOE employs a Mixture of Low-Rank Adaptation (MoLoRA) technique, allowing for efficient parameter usage by maintaining base model parameters static while adapting through a minimal set of trainable parameters. We demonstrate that MING-MOE achieves state-of-the-art (SOTA) performance on over 20 medical tasks, illustrating a significant improvement over existing models. This approach not only extends the capabilities of medical language models but also improves inference efficiency.

📄 PDF Abstract BibTeX arXiv:2404.09027

Code (2)

mediabrain-sjtu/ming 공식 구현 pytorch
mediabrain-sjtu/medicalgpt-zh pytorch

Tasks

DiversityLanguage ModelingLanguage ModellingLarge Language ModelMulti-Task LearningNatural Language Understanding

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
BASE 설명 없음

Similar Papers 제목 키워드 기반

Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

2025-09-11 · Qiuhui Chen, Xuancheng Yao, Huping Ye, Yi Hong arxiv

Understanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent adv…

Self-Supervised LearningRepresentation Learning

Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA

2026-06-18 · Yuetian Du, Yucheng Wang, Ming Kong, Tian Liang 외 arxiv

Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual accuracy, potentially leading to misdiagnosis or overlooking correct advice. This …

Visual Question Answering

BMRetriever: Tuning Large Language Models as Better Biomedical Text Retrievers

2024-04-29 · ran Xu, Wenqi Shi, Yue Yu, Yuchen Zhuang 외

Developing effective biomedical retrieval models is important for excelling at knowledge-intensive biomedical tasks but still challenging due to the deficiency of sufficient publicly annotated biomedical data and computa…

RetrievalUnsupervised Pre-training

The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models

2024-02-12 · Ayo Adedeji, Sarita Joshi, Brendan Doohan

In the rapidly evolving landscape of medical documentation, transcribing clinical dialogues accurately is increasingly paramount. This study explores the potential of Large Language Models (LLMs) to enhance the accuracy …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Semantic Textual Similarityspeaker-diarization+3

MiniGPT-Med: Large Language Model as a General Interface for Radiology Diagnosis

2024-07-04 · Asma Alkhaldi, Raneem Alnajim, Layan Alabdullatef, Rawan Alyahya 외

Recent advancements in artificial intelligence (AI) have precipitated significant breakthroughs in healthcare, particularly in refining diagnostic procedures. However, previous studies have often been constrained to limi…

DiagnosticLanguage ModelingLanguage ModellingLarge Language Model+4