paper-with-me

Papers

Understanding Task Aggregation for Generalizable Ultrasound Foundation Models

2026-03-18 · Fangyijie Wang, Tanya Akumu, Vien Ngoc Dang, Amelia Jiménez-Sánchez, Jieyun Bai, Guénolé Silvestre, Karim Lekadir, Kathleen M. Curran arxiv

Foundation models promise to unify multiple clinical tasks within a single framework, but recent ultrasound studies report that unified models can underperform task-specific baselines. We hypothesize that this degradation arises not from model capacity limitations, but from task aggregation strategies that ignore interactions between task heterogeneity and available training data scale. In this work, we systematically analyze when heterogeneous ultrasound tasks can be jointly learned without performance loss, establishing practical criteria for task aggregation in unified clinical imaging models. We introduce M2DINO, a multi-organ, multi-task framework built on DINOv3 with task-conditioned Mixture-of-Experts blocks for adaptive capacity allocation. We systematically evaluate 27 ultrasound tasks spanning segmentation, classification, detection, and regression under three paradigms: task-specific, clinically-grouped, and all-task unified training. Our results show that aggregation effectiveness depends strongly on training data scale. While clinically-grouped training can improve performance in data-rich settings, it may induce substantial negative transfer in low-data settings. In contrast, all-task unified training exhibits more consistent performance across clinical groups. We further observe that task sensitivity varies by task type in our experiments: segmentation shows the largest performance drops compared with regression and classification. These findings provide practical guidance for ultrasound foundation models, emphasizing that aggregation strategies should jointly consider training data availability and task characteristics rather than relying on clinical taxonomy alone.

📄 PDF Abstract BibTeX arXiv:2603.18123

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Fully Open and Generalizable Foundation Model for Ultrasound Clinical Applications

2025-09-15 · Hongyuan Zhang, Yuheng Wu, Mingyang Zhao, Zhiwei Chen 외 arxiv

Artificial intelligence (AI) that can effectively learn ultrasound representations by integrating multi-source data holds significant promise for advancing clinical care. However, the scarcity of large labeled datasets i…

Self-Supervised LearningLesion Segmentation

Baseline Method of the Foundation Model Challenge for Ultrasound Image Analysis

2026-02-01 · Bo Deng, Yitong Tang, Jiake Li, Yuxin Huang 외 arxiv

Ultrasound (US) imaging exhibits substantial heterogeneity across anatomical structures and acquisition protocols, posing significant challenges to the development of generalizable analysis models. Most existing methods …

Multi-Task Learning

OpenUS: A Fully Open-Source Foundation Model for Ultrasound Image Analysis via Self-Adaptive Masked Contrastive Learning

2025-11-14 · Xiaoyu Zheng, Xu Chen, Awais Rauf, Qifan Fu 외 arxiv

Ultrasound (US) is one of the most widely used medical imaging modalities, thanks to its low cost, portability, real-time feedback, and absence of ionizing radiation. However, US image interpretation remains highly opera…

Contrastive Learning

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

2025-11-27 · Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao 외 arxiv

Despite significant strides in medical foundation models, the ultrasound domain lacks a comprehensive solution capable of bridging low-level Ultrasound Grounded Perception (e.g., segmentation, localization) and high-leve…

FeatureNeRF: Learning Generalizable NeRFs by Distilling Foundation Models

2023-03-22 · ICCV 2023 1 · Jianglong Ye, Naiyan Wang, Xiaolong Wang

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semanti…

NeRFNeural RenderingNovel View Synthesis