Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs
Recent advances in multimodal ECG representation learning center on aligning ECG signals with paired free-text reports. However, suboptimal alignment persists due to the complexity of medical language and the reliance on a full 12-lead setup, which is often unavailable in under-resourced settings. To tackle these issues, we propose K-MERL, a knowledge-enhanced multimodal ECG representation learning framework. K-MERL leverages large language models to extract structured knowledge from free-text reports and employs a lead-aware ECG encoder with dynamic lead masking to accommodate arbitrary lead inputs. Evaluations on six external ECG datasets show that K-MERL achieves state-of-the-art performance in zero-shot classification and linear probing tasks, while delivering an average 16% AUC improvement over existing methods in partial-lead zero-shot classification.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation Learningzero-shot-classificationZero-Shot LearningSimilar Papers 제목 키워드 기반
VideoAdviser: Video Knowledge Distillation for Multimodal Transfer Learning
Multimodal transfer learning aims to transform pretrained representations of diverse modalities into a common domain space for effective multimodal fusion. However, conventional systems are typically built on the assumpt…
Knowledge DistillationregressionSentiment AnalysisTransfer LearningKBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution
The rapid progress of multimodal large language models (MLLMs) calls for more reliable evaluation protocols. Existing static benchmarks suffer from the potential risk of data contamination and saturation, leading to infl…
Otter-Knowledge: benchmarks of multimodal knowledge graph representation learning from different sources for drug discovery
Recent research on predicting the binding affinity between drug molecules and proteins use representations learned, through unsupervised learning techniques, from large databases of molecule SMILES and protein sequences.…
Drug DiscoveryGraph Representation LearningKnowledge GraphsPrediction+1Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
Multimodal Large Language Models (MLLMs) have propelled the field of few-shot action recognition (FSAR). However, preliminary explorations in this area primarily focus on generating captions to form a suboptimal feature-…
Action RecognitionMetric LearningMultimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
Text response generation for multimodal task-oriented dialog systems, which aims to generate the proper text response given the multimodal context, is an essential yet challenging task. Although existing efforts have ach…
DecoderLanguage ModelingLanguage ModellingResponse Generation