paper-with-me

Papers

Benchmarking ECG FMs: A Reality Check Across Clinical Tasks

2025-09-29 · M A Al-Masud, Juan Miguel Lopez Alcaraz, Nils Strodthoff arxiv

The 12-lead electrocardiogram (ECG) is a long-standing diagnostic tool. Yet machine learning for ECG interpretation remains fragmented, often limited to narrow tasks or datasets. FMs promise broader adaptability, but fundamental questions remain: Which architectures generalize best? How do models scale with limited labels? What explains performance differences across model families? We benchmarked eight ECG FMs on 26 clinically relevant tasks using 12 public datasets comprising 1,650 regression and classification targets. Models were evaluated under fine-tuning and frozen settings, with scaling analyses across dataset sizes. Results show heterogeneous performance across domains: in adult ECG interpretation, three FMs consistently outperformed strong supervised baselines. In contrast, ECG-CPC, a compact structured state-space model, dominated 5 of 7 task categories, demonstrating that architecture matters more than scale. FMs improved label efficiency 3.3-9x over supervised baselines, though scaling behaviors varied across architectures. Representation analysis reveals that models with similar performance learn markedly different internal structures, suggesting multiple viable paths to effective ECG representation. Overall, while FMs show promise for adult ECG analysis, substantial gaps remain in cardiac structure, outcome prediction, and patient characterization. ECG-CPC's strong performance despite being orders of magnitude smaller challenges the assumption that FM quality requires massive scale, highlighting architectural inductive biases as an untapped opportunity.

📄 PDF Abstract BibTeX arXiv:2509.25095

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Failure Detection in Medical Image Classification: A Reality Check and Benchmarking Testbed

2022-05-27 · Melanie Bernhardt, Fabio De Sousa Ribeiro, Ben Glocker

Failure detection in automated image classification is a critical safeguard for clinical deployment. Detected failure cases can be referred to human assessment, ensuring patient safety in computer-aided clinical decision…

BenchmarkingBinary ClassificationDecision MakingGeneral Classification+4

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI

2024-12-17 · Deep Bhatt, Surya Ayyagari, Anuruddh Mishra

Diagnostic errors in healthcare persist as a critical challenge, with increasing numbers of patients turning to online resources for health information. While AI-powered healthcare chatbots show promise, there exists no …

BenchmarkingChatbotDiagnostic

GOLDMARK: Governed Outcome-Linked Diagnostic Model Assessment Reference Kit

2026-03-21 · Chad Vanderbilt, Gabriele Campanella, Siddharth Singi, Swaraj Nanda 외 arxiv

Computational biomarkers (CBs) are histopathology-derived patterns extracted from hematoxylin-eosin (H&E) whole-slide images (WSIs) using artificial intelligence (AI) to predict therapeutic response or prognosis. Recentl…

Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models

2026-01-11 · Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng 외 arxiv

Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challen…

OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs

2024-05-09 · Yuxia Wang, Minghan Wang, Hasan Iqbal, Georgi Georgiev 외

The increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the factual accuracy of their outputs. Difficulties lie in assessing the factuality of free-for…

BenchmarkingFact Checking