paper-with-me

홈 › Papers

TotalFM: An Organ-Separated 3D-CT Foundation Model Leveraging Large-Scale Routine Clinical Radiology Data

2026-01-01 · Kohei Yamamoto, Tomohiro Kikuchi arxiv

While foundation models in radiology are expected to be applied to various clinical tasks, computational cost constraints remain a major challenge when training on 3D-CT volumetric data. In this study, we propose TotalFM, a radiological foundation model that efficiently learns the correspondence between 3D-CT images and linguistic expressions based on the concept of organ separation, utilizing a large-scale dataset of 140,000 series. By automating the creation of organ volume and finding-sentence pairs through segmentation techniques and Large Language Model (LLM)-based radiology report processing, and by combining self-supervised pre-training via VideoMAE with contrastive learning using volume-text pairs, we aimed to balance computational efficiency and representation capability. In zero-shot organ-wise lesion classification tasks, the proposed model achieved higher F1 scores in 83% (5/6) of organs compared to CT-CLIP and 64% (9/14) of organs compared to Merlin. These results suggest that the proposed model exhibits high generalization performance in a clinical evaluation setting using actual radiology report sentences. Furthermore, in zero-shot finding-wise lesion classification tasks, our model achieved a higher AUROC in 83% (25/30) of finding categories compared to Merlin. We also confirmed performance comparable to existing Vision-Language Models (VLMs) in radiology report generation tasks. Our results demonstrate that the organ-separated learning framework can serve as a realistic and effective design guideline for the practical implementation of 3D-CT foundation models. The source code and pretrained models are publicly available at https://github.com/jichi-labo/TotalFM.

📄 PDF Abstract BibTeX arXiv:2601.00260

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyContrastive Learning

Similar Papers 제목 키워드 기반

Learning task-specific subspaces via interventional post-training of speech foundation models

2026-06-16 · Jack Cox, Jon Barker arxiv

Speech foundation models, pre-trained on large corpora of unlabelled speech data, produce general-purpose representations which are useful across tasks. However, these representations encode information about salient spe…

Contrastive LearningSpeaker VerificationKeyword Spotting

OrganLens: Organ-Specific Representation Learning for CT Foundation Models

2026-07-28 · Zhixuan Ge, Anqi Li, Sadeer Al-Kindi, Hanwen Xu 외 arxiv

A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ with…

Representation LearningImage-to-Text Retrieval

Confederated Machine Learning on Horizontally and Vertically Separated Medical Data for Large-Scale Health System Intelligence

2019-10-04 · ICLR 2020 1 · Dianbo Liu, Kathe Fox, Griffin Weber, Tim Miller

Health information is generally fragmented across silos. Though it is technically feasible to unite data for analysis in a manner that underpins a rapid learning healthcare system, privacy concerns and regulatory barrier…

BIG-bench Machine LearningFederated Learning

Towards Foundation Models and Few-Shot Parameter-Efficient Fine-Tuning for Volumetric Organ Segmentation

2023-03-29 · Julio Silva-Rodríguez, Jose Dolz, Ismail Ben Ayed

The recent popularity of foundation models and the pre-train-and-adapt paradigm, where a large-scale model is transferred to downstream tasks, is gaining attention for volumetric medical image segmentation. However, curr…

Image SegmentationMedical Image SegmentationOrgan Segmentationparameter-efficient fine-tuning+4

BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning

2025-05-29 · Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo, Matthew J Thompson 외

Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale co…

Contrastive Learning