paper-with-me

Papers

PRISM: Synergizing Vision Foundation Models via Self-organized Expert Specialization

2026-06-02 · Ying Tang, Dong Li, Youjia Zhang, Zikai Song, Junqing Yu, Wei Yang arxiv

Unifying the complementary strengths of diverse Vision Foundation Models (VFMs) into a single efficient model is highly desirable but challenged by the negative transfer inherent in monolithic distillation. To address these feature conflicts, we introduce \textbf{PRISM}, a novel dual-stream Mixture-of-Experts (MoE) framework that synergizes VFMs via modular specialization. We propose a two-stage paradigm: (1) expertise deconstruction, where a teacher-conditional router guides experts to specialize in distinct representational subspaces to mitigate interference, followed by (2) dynamic recomposition, where the router learns to assemble these experts into tailored computational pathways for downstream tasks. Experiments on PASCAL-Context and NYUD-v2 show that \textbf{PRISM} establishes a new state of the art, validating that sparse, emergent specialization is a scalable approach for integrating diverse visual knowledge.

📄 PDF Abstract BibTeX arXiv:2606.03444

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PRISM-CTG: A Foundation Model for Cardiotocography Analysis with Multi-View SSL

2026-04-09 · Sheng Wong, Ravi Shankar, Beth Albert, Hao Fei 외 arxiv

Supervised deep learning models for automated CTG analysis are typically constrained by narrowly curated labelled datasets and limited patient cohorts, leaving substantial volumes of physiologically informative clinical …

Representation Learning

MOOZY: A Patient-First Foundation Model for Computational Pathology

2026-03-27 · Yousef Kotp, Vincent Quoc-Huy Trinh, Christopher Pal, Mahdi S. Hosseini arxiv

Computational pathology needs whole-slide image (WSI) foundation models that transfer across diverse clinical tasks, yet current approaches remain largely slide-centric, often depend on private data and expensive paired-…

Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task Learning

2024-10-18 · Yuxiang Lu, Shengcao Cao, Yu-Xiong Wang

Vision Foundation Models (VFMs) have demonstrated outstanding performance on numerous downstream tasks. However, due to their inherent representation biases originating from different training paradigms, VFMs exhibit adv…

Multi-Task LearningTransfer Learning

PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue

2025-06-16 · George Shaikovski, Eugene Vorontsov, Adam Casson, Julian Viret 외

Recent pathology foundation models can provide rich tile-level representations but fall short of delivering general-purpose clinical utility without further extensive model development. These models lack whole-slide imag…

DiagnosticLanguage ModelingLanguage Modelling

PRISM: Distributed Inference for Foundation Models at Edge

2025-07-16 · Muhammad Azlan Qazi, Alexandros Iosifidis, Qi Zhang arxiv

Foundation models (FMs) have achieved remarkable success across a wide range of applications, from image classification to natural langurage processing, but pose significant challenges for deployment at edge. This has sp…

Image Classification