MedPix 2.0: A Comprehensive Multimodal Biomedical Data set for Advanced AI Applications
The increasing interest in developing Artificial Intelligence applications in the medical domain, suffers from the lack of high-quality data set, mainly due to privacy-related issues. Moreover, the recent rising of Large Multimodal Models (LMM) leads to a need for multimodal medical data sets, where clinical reports and findings are attached to the corresponding CT or MR scans. This paper illustrates the entire workflow for building the data set MedPix 2.0. Starting from the well-known multimodal data set MedPix, mainly used by physicians, nurses and healthcare students for Continuing Medical Education purposes, a semi-automatic pipeline was developed to extract visual and textual data followed by a manual curing procedure where noisy samples were removed, thus creating a MongoDB database. Along with the data set, we developed a GUI aimed at navigating efficiently the MongoDB instance, and obtaining the raw data that can be easily used for training and/or fine-tuning LMMs. To enforce this point, we also propose a CLIP-based model trained on MedPix 2.0 for scanning modality and location classification tasks. MedPix 2.0 is available on GitHub
Code (1)
Tasks
Knowledge GraphsRAGMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey
Explainable artificial intelligence (XAI) has become increasingly important in biomedical image analysis to promote transparency, trust, and clinical adoption of DL models. While several surveys have reviewed XAI techniq…
Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)SurveyBMMDetect: A Multimodal Deep Learning Framework for Comprehensive Biomedical Misconduct Detection
Academic misconduct detection in biomedical research remains challenging due to algorithmic narrowness in existing methods and fragmented analytical pipelines. We present BMMDetect, a multimodal deep learning framework t…
ArticlesFeature ImportanceMultimodal Deep LearningA Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications
Deep learning has revolutionized biomedical research by providing sophisticated methods to handle complex, high-dimensional data. Multimodal deep learning (MDL) further enhances this capability by integrating diverse dat…
Deep LearningMultimodal Deep LearningMulti-modal Pre-training for Medical Vision-language Understanding and Generation: An Empirical Study with A New Benchmark
With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective f…
Image-text RetrievalMedical Report GenerationQuestion AnsweringRetrieval+2BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
Biomedical data is inherently multimodal, comprising physical measurements and natural language narratives. A generalist biomedical AI model needs to simultaneously process different modalities of data, including text an…
ArticlesMedical Visual Question AnsweringPneumonia DetectionQuestion Answering+3