paper-with-me

홈 › Papers

Shazam: Unifying Multiple Foundation Models for Advanced Computational Pathology

2025-03-02 · Wenhui Lei, Anqi Li, Yusheng Tan, HanYu Chen, Xiaofan Zhang

Foundation Models (FMs) in computational pathology (CPath) have significantly advanced the extraction of meaningful features from histopathology image datasets, achieving strong performance across various clinical tasks. Despite their impressive performance, these models often exhibit variability when applied to different tasks, prompting the need for a unified framework capable of consistently excelling across various applications. In this work, we propose Shazam, a novel framework designed to efficiently combine multiple CPath models. Unlike previous approaches that train a fixed-parameter FM, Shazam dynamically extracts and refines information from diverse FMs for each specific task. To ensure that each FM contributes effectively without dominance, a novel distillation strategy is applied, guiding the student model with features from all teacher models, which enhances its generalization ability. Experimental results on two pathology patch classification datasets demonstrate that Shazam outperforms existing CPath models and other fusion methods. Its lightweight, flexible design makes it a promising solution for improving CPath analysis in real-world settings. Code will be available at https://github.com/Tuner12/Shazam.

📄 PDF Abstract BibTeX arXiv:2503.00736

Code (1)

tuner12/shazam 공식 구현 pytorch

Similar Papers 제목 키워드 기반

SHAZAM: Self-Supervised Change Monitoring for Hazard Detection and Mapping

2025-03-01 · Samuel Garske, Konrad Heidler, Bradley Evans, KC Wong 외

The increasing frequency of environmental hazards due to climate change underscores the urgent need for effective monitoring systems. Current approaches either rely on expensive labelled datasets, struggle with seasonal …

Advancing Audio Fingerprinting Accuracy Addressing Background Noise and Distortion Challenges

2024-02-21 · Navin Kamuni, Sathishkumar Chintala, Naveen Kunchakuri, Jyothi Swaroop Arlagadda Narasimharaju 외

Audio fingerprinting, exemplified by pioneers like Shazam, has transformed digital audio recognition. However, existing systems struggle with accuracy in challenging conditions, limiting broad applicability. This researc…

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

2024-10-29 · Shenghao Xie, Wenqiang Zu, Mingyang Zhao, Duo Su 외

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to v…

Survey

On the Conic Complementarity of Planar Contacts

2025-09-30 · Yann de Mont-Marin, Louis Montaut, Jean Ponce, Martial Hebert 외 arxiv

We present a unifying theoretical result that connects two foundational principles in robotics: the Signorini law for point contacts, which underpins many simulation methods for preventing object interpenetration, and th…

Vision-Centric Activation and Coordination for Multimodal Large Language Models

2025-10-16 · Yunnan Wang, Fan Lu, Kecheng Zheng, Ziyuan Huang 외 arxiv

Multimodal large language models (MLLMs) integrate image features from visual encoders with LLMs, demonstrating advanced comprehension capabilities. However, mainstream MLLMs are solely supervised by the next-token predi…