paper-with-me

홈 › Papers

MedPix 2.0: A Comprehensive Multimodal Biomedical Data set for Advanced AI Applications

2024-07-03 · Irene Siragusa, Salvatore Contino, Massimo La Ciura, Rosario Alicata, Roberto Pirrone

The increasing interest in developing Artificial Intelligence applications in the medical domain, suffers from the lack of high-quality data set, mainly due to privacy-related issues. Moreover, the recent rising of Large Multimodal Models (LMM) leads to a need for multimodal medical data sets, where clinical reports and findings are attached to the corresponding CT or MR scans. This paper illustrates the entire workflow for building the data set MedPix 2.0. Starting from the well-known multimodal data set MedPix, mainly used by physicians, nurses and healthcare students for Continuing Medical Education purposes, a semi-automatic pipeline was developed to extract visual and textual data followed by a manual curing procedure where noisy samples were removed, thus creating a MongoDB database. Along with the data set, we developed a GUI aimed at navigating efficiently the MongoDB instance, and obtaining the raw data that can be easily used for training and/or fine-tuning LMMs. To enforce this point, we also propose a CLIP-based model trained on MedPix 2.0 for scanning modality and location classification tasks. MedPix 2.0 is available on GitHub

📄 PDF Abstract BibTeX arXiv:2407.02994

Code (1)

chilab1/medpix-2.0 공식 구현 pytorch

Tasks

Knowledge GraphsRAG

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Explainable Artificial Intelligence in Biomedical Image Analysis: A Comprehensive Survey

2025-07-09 · Getamesay Haile Dagnaw, Yanming Zhu, Muhammad Hassan Maqsood, Wencheng Yang 외

Explainable artificial intelligence (XAI) has become increasingly important in biomedical image analysis to promote transparency, trust, and clinical adoption of DL models. While several surveys have reviewed XAI techniq…

Explainable artificial intelligenceExplainable Artificial Intelligence (XAI)Survey

BMMDetect: A Multimodal Deep Learning Framework for Comprehensive Biomedical Misconduct Detection

2025-05-09 · Yize Zhou, Jie Zhang, Meijie Wang, Lun Yu

Academic misconduct detection in biomedical research remains challenging due to algorithmic narrowness in existing methods and fragmented analytical pipelines. We present BMMDetect, a multimodal deep learning framework t…

ArticlesFeature ImportanceMultimodal Deep Learning

A Systematic Review of Intermediate Fusion in Multimodal Deep Learning for Biomedical Applications

2024-08-02 · Valerio Guarrasi, Fatih Aksu, Camillo Maria Caruso, Francesco Di Feola 외

Deep learning has revolutionized biomedical research by providing sophisticated methods to handle complex, high-dimensional data. Multimodal deep learning (MDL) further enhances this capability by integrating diverse dat…

Deep LearningMultimodal Deep Learning

Multi-modal Pre-training for Medical Vision-language Understanding and Generation: An Empirical Study with A New Benchmark

2023-06-10 · Li Xu, Bo Liu, Ameer Hamza Khan, Lu Fan 외

With the availability of large-scale, comprehensive, and general-purpose vision-language (VL) datasets such as MSCOCO, vision-language pre-training (VLP) has become an active area of research and proven to be effective f…

Image-text RetrievalMedical Report GenerationQuestion AnsweringRetrieval+2

BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

2023-03-02 · Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu 외

Biomedical data is inherently multimodal, comprising physical measurements and natural language narratives. A generalist biomedical AI model needs to simultaneously process different modalities of data, including text an…

ArticlesMedical Visual Question AnsweringPneumonia DetectionQuestion Answering+3