paper-with-me

Papers

A Medical Data-Effective Learning Benchmark for Highly Efficient Pre-training of Foundation Models

2024-01-31 · Wenxuan Yang, Weimin Tan, Yuqi Sun, Bo Yan

Foundation models, pre-trained on massive datasets, have achieved unprecedented generalizability. However, is it truly necessary to involve such vast amounts of data in pre-training, consuming extensive computational resources? This paper introduces data-effective learning, aiming to use data in the most impactful way to pre-train foundation models. This involves strategies that focus on data quality rather than quantity, ensuring the data used for training has high informational value. Data-effective learning plays a profound role in accelerating foundation model training, reducing computational costs, and saving data storage, which is very important as the volume of medical data in recent years has grown beyond many people's expectations. However, due to the lack of standards and comprehensive benchmarks, research on medical data-effective learning is poorly studied. To address this gap, our paper introduces a comprehensive benchmark specifically for evaluating data-effective learning in the medical field. This benchmark includes a dataset with millions of data samples from 31 medical centers (DataDEL), a baseline method for comparison (MedDEL), and a new evaluation metric (NormDEL) to objectively measure data-effective learning performance. Our extensive experimental results show the baseline MedDEL can achieve performance comparable to the original large dataset with only 5% of the data. Establishing such an open data-effective learning benchmark is crucial for the medical foundation model research community because it facilitates efficient data use, promotes collaborative breakthroughs, and fosters the development of cost-effective, scalable, and impactful healthcare solutions.

📄 PDF Abstract BibTeX arXiv:2401.17542

Code (1)

shadow2469/data-effective-learning-a-comprehensive-medical-benchmark 공식 구현

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

BioAug: Conditional Generation based Data Augmentation for Low-Resource Biomedical NER

2023-05-18 · Sreyan Ghosh, Utkarsh Tyagi, Sonal Kumar, Dinesh Manocha

Biomedical Named Entity Recognition (BioNER) is the fundamental task of identifying named entities from biomedical text. However, BioNER suffers from severe data scarcity and lacks high-quality labeled data due to the hi…

Data Augmentationnamed-entity-recognitionNamed Entity RecognitionNER

Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation

2025-11-12 · Frank Li, Theo Dapamede, Mohammadreza Chavoshi, Young Seok Jeon 외 arxiv

Foundation models (FMs) promise to generalize medical imaging, but their effectiveness varies. It remains unclear how pre-training domain (medical vs. general), paradigm (e.g., text-guided), and architecture influence em…

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

2026-06-02 · Henry He, Johann Frei, Raphael Schmitt arxiv

Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limited by older architectures or restricted training data. We present Ch…

Medical Named Entity RecognitionText ClassificationDomain Adaptation

Less Is More: A Comparison of Active Learning Strategies for 3D Medical Image Segmentation

2022-07-02 · Josafat-Mattias Burmeister, Marcel Fernandez Rosas, Johannes Hagemann, Jonas Kordt 외

Since labeling medical image data is a costly and labor-intensive process, active learning has gained much popularity in the medical image segmentation domain in recent years. A variety of active learning strategies have…

Active LearningBenchmarkingImage SegmentationMedical Image Segmentation+2

Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval

2025-01-15 · Demetrio Deanda, Yuktha Priya Masupalli, Jeong Yang, Young Lee 외

Medical images and reports offer invaluable insights into patient health. The heterogeneity and complexity of these data hinder effective analysis. To bridge this gap, we investigate contrastive learning models for cross…

BenchmarkingContrastive LearningRetrieval