paper-with-me

Papers

Multimodal Foundation Models For Echocardiogram Interpretation

2023-08-29 · Matthew Christensen, Milos Vukadinovic, Neal Yuan, David Ouyang

Multimodal deep learning foundation models can learn the relationship between images and text. In the context of medical imaging, mapping images to language concepts reflects the clinical task of diagnostic image interpretation, however current general-purpose foundation models do not perform well in this context because their training corpus have limited medical text and images. To address this challenge and account for the range of cardiac physiology, we leverage 1,032,975 cardiac ultrasound videos and corresponding expert interpretations to develop EchoCLIP, a multimodal foundation model for echocardiography. EchoCLIP displays strong zero-shot (not explicitly trained) performance in cardiac function assessment (external validation left ventricular ejection fraction mean absolute error (MAE) of 7.1%) and identification of implanted intracardiac devices (areas under the curve (AUC) between 0.84 and 0.98 for pacemakers and artificial heart valves). We also developed a long-context variant (EchoCLIP-R) with a custom echocardiography report text tokenizer which can accurately identify unique patients across multiple videos (AUC of 0.86), identify clinical changes such as orthotopic heart transplants (AUC of 0.79) or cardiac surgery (AUC 0.77), and enable robust image-to-text search (mean cross-modal retrieval rank in the top 1% of candidate text reports). These emergent capabilities can be used for preliminary assessment and summarization of echocardiographic findings.

📄 PDF Abstract BibTeX arXiv:2308.15670

Code (1)

echonet/echo_CLIP 공식 구현 pytorch

Tasks

Cross-Modal RetrievalDiagnosticImage to textMultimodal Deep LearningRetrieval

Similar Papers 제목 키워드 기반

Semi-Supervised Multimodal Multi-Instance Learning for Aortic Stenosis Diagnosis

2024-03-09 · Zhe Huang, Xiaowei Yu, Benjamin S. Wessler, Michael C. Hughes

Automated interpretation of ultrasound imaging of the heart (echocardiograms) could improve the detection and treatment of aortic stenosis (AS), a deadly heart disease. However, existing deep learning pipelines for asses…

Multiple Instance Learning

MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management

2026-03-23 · Jack W O'Sullivan, Mohammad Asadi, Lennart Elbe, Akshay Chaudhari 외 arxiv

Cardiovascular disease remains the leading cause of global mortality, with progress hindered by human interpretation of complex cardiac tests. Current AI vision-language models are limited to single-modality inputs and a…

Fast and accurate classification of echocardiograms using deep learning

2017-06-27 · Ali Madani, Ramy Arnaout, Mohammad Mofrad, Rima Arnaout

Echocardiography is essential to modern cardiology. However, human interpretation limits high throughput analysis, limiting echocardiography from reaching its full clinical and research potential for precision medicine. …

ClassificationDeep LearningGeneral ClassificationOverall - Test

Echocardiogram Foundation Model -- Application 1: Estimating Ejection Fraction

2023-11-21 · Adil Dahlan, Cyril Zakka, Abhinav Kumar, Laura Tang 외

Cardiovascular diseases stand as the primary global cause of mortality. Among the various imaging techniques available for visualising the heart and evaluating its function, echocardiograms emerge as the preferred choice…

modelSelf-Supervised Learning

EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation

2024-10-13 · Milos Vukadinovic, Xiu Tang, Neal Yuan, Paul Cheng 외

Echocardiography is the most widely used cardiac imaging modality, capturing ultrasound video data to assess cardiac structure and function. Artificial intelligence (AI) in echocardiography has the potential to streamlin…

Contrastive LearningLanguage ModelingLanguage Modelling