paper-with-me

홈 › Papers

MedDChest: A Content-Aware Multimodal Foundational Vision Model for Thoracic Imaging

2025-11-06 · Mahmoud Soliman, Islam Osman, Mohamed S. Shehata, Rasika Rajapakshe arxiv

The performance of vision models in medical imaging is often hindered by the prevailing paradigm of fine-tuning backbones pre-trained on out-of-domain natural images. To address this fundamental domain gap, we propose MedDChest, a new foundational Vision Transformer (ViT) model optimized specifically for thoracic imaging. We pre-trained MedDChest from scratch on a massive, curated, multimodal dataset of over 1.2 million images, encompassing different modalities including Chest X-ray and Computed Tomography (CT) compiled from 10 public sources. A core technical contribution of our work is Guided Random Resized Crops, a novel content-aware data augmentation strategy that biases sampling towards anatomically relevant regions, overcoming the inefficiency of standard cropping techniques on medical scans. We validate our model's effectiveness by fine-tuning it on a diverse set of downstream diagnostic tasks. Comprehensive experiments empirically demonstrate that MedDChest significantly outperforms strong, publicly available ImageNet-pretrained models. By establishing the superiority of large-scale, in-domain pre-training combined with domain-specific data augmentation, MedDChest provides a powerful and robust feature extractor that serves as a significantly better starting point for a wide array of thoracic diagnostic tasks. The model weights will be made publicly available to foster future research and applications.

📄 PDF Abstract BibTeX arXiv:2511.04016

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

Needle In A Multimodal Haystack

2024-06-11 · Weiyun Wang, Shuibo Zhang, Yiming Ren, Yuchen Duan 외

With the rapid advancement of multimodal large language models (MLLMs), their evaluation has become increasingly comprehensive. However, understanding long multimodal content, as a foundational ability for real-world app…

Retrieval

MMCFND: Multimodal Multilingual Caption-aware Fake News Detection for Low-resource Indic Languages

2024-10-14 · Shubhi Bansal, Nishit Sushil Singh, Shahid Shafi Dar, Nagendra Kumar

The widespread dissemination of false information through manipulative tactics that combine deceptive text and images threatens the integrity of reliable sources of information. While there has been research on detecting…

ArticlesDescriptiveFake News DetectionImage Captioning

Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs

2025-02-16 · Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang 외

Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. However, ensuring the safety of these models remains a signific…

Benchmarking

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

2025-01-01 · Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun 외

Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges l…

Optical Character Recognition (OCR)

Hyperbolic Safety-Aware Vision-Language Models

2025-03-15 · CVPR 2025 1 · Tobia Poppi, Tejaswi Kasarla, Pascal Mettes, Lorenzo Baraldi 외

Addressing the retrieval of unsafe content from vision-language models such as CLIP is an important step towards real-world integration. Current efforts have relied on unlearning techniques that try to erase the model's …