paper-with-me

홈 › Papers

ProFound: A moderate-sized vision foundation model for multi-task prostate imaging

2026-03-04 · Yipei Wang, Yinsong Xu, Weixi Yi, Shaheer Ullah Saeed, Natasha Thorley, Alexander Ng, Yukun Zhou, Wen Yan, Dean Barratt, Shonit Punwani, Veeru Kasivisvanathan, Mark Emberton, Daniel C. Alexander, Yipeng Hu arxiv

Many diagnostic and therapeutic clinical tasks for prostate cancer increasingly rely on multi-parametric MRI. Automating these tasks is challenging because they necessitate expert interpretations, which are difficult to scale to capitalise on modern deep learning. Although modern automated systems achieve expert-level performance in isolated tasks, their general clinical utility remains limited by the requirement of large task-specific labelled datasets. In this paper, we present ProFound, a domain-specialised vision foundation model for volumetric prostate mpMRI. ProFound is pre-trained using several variants of self-supervised approaches on a diverse, multi-institutional collection of 5,000 patients, with a total of over 22,000 unique 3D MRI volumes (over 1,800,000 2D image slices). We conducted a systematic evaluation of ProFound across a broad spectrum of $11$ downstream clinical tasks on over 3,000 independent patients, including prostate cancer detection, Gleason grading, lesion localisation, gland volume estimation, zonal and surrounding structure segmentation. Experimental results demonstrate that finetuned ProFound consistently outperforms or remains competitive with state-of-the-art specialised models and existing medical vision foundation models trained/finetuned on the same data.

📄 PDF Abstract BibTeX arXiv:2603.03961

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ConvNets Match Vision Transformers at Scale

2023-10-25 · Samuel L. Smith, Andrew Brock, Leonard Berrada, Soham De

Many researchers believe that ConvNets perform well on small or moderately sized datasets, but are not competitive with Vision Transformers when given access to datasets on the web-scale. We challenge this belief by eval…

LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model

2024-03-29 · Musashi Hinck, Matthew L. Olson, David Cobbley, Shao-Yen Tseng 외

We train a suite of multimodal foundation models (MMFM) using the popular LLaVA framework with the recently released Gemma family of large language models (LLMs). Of particular interest is the 2B parameter Gemma model, w…

Language ModelingLanguage Modelling

Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

2025-04-11 · Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao 외

This technical report presents a cost-efficient strategy for training a video generation foundation model. We present a mid-sized research model with approximately 7 billion parameters (7B) called Seaweed-7B trained from…

GPUVideo Generation

Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model

2023-10-18 · Qi Jia, Siyu Ren, Yizhu Liu, Kenny Q. Zhu

Despite tremendous improvements in natural language generation, summarization models still suffer from the unfaithfulness issue. Previous work evaluates faithfulness either using models trained on the other tasks or in-d…

Language ModelingLanguage ModellingText GenerationText Summarization

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

2026-04-17 · Dian Shao, Zhengzheng Xu, Peiyang Wang, Like Liu 외 arxiv

UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi-step instructions over long horizons. Existing zero-shot methods r…

Vision-Language Navigation