paper-with-me

홈 › Papers

Fine-Grained ECG-Text Contrastive Learning via Waveform Understanding Enhancement

2025-05-17 · Haitao Li, Che Liu, Zhengyao Ding, Ziyi Liu, Zhengxing Huang

Electrocardiograms (ECGs) are essential for diagnosing cardiovascular diseases. While previous ECG-text contrastive learning methods have shown promising results, they often overlook the incompleteness of the reports. Given an ECG, the report is generated by first identifying key waveform features and then inferring the final diagnosis through these features. Despite their importance, these waveform features are often not recorded in the report as intermediate results. Aligning ECGs with such incomplete reports impedes the model's ability to capture the ECG's waveform features and limits its understanding of diagnostic reasoning based on those features. To address this, we propose FG-CLEP (Fine-Grained Contrastive Language ECG Pre-training), which aims to recover these waveform features from incomplete reports with the help of large language models (LLMs), under the challenges of hallucinations and the non-bijective relationship between waveform features and diagnoses. Additionally, considering the frequent false negatives due to the prevalence of common diagnoses in ECGs, we introduce a semantic similarity matrix to guide contrastive learning. Furthermore, we adopt a sigmoid-based loss function to accommodate the multi-label nature of ECG-related tasks. Experiments on six datasets demonstrate that FG-CLEP outperforms state-of-the-art methods in both zero-shot prediction and linear probing across these datasets.

📄 PDF Abstract BibTeX arXiv:2505.11939

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningDiagnosticSemantic SimilaritySemantic Textual Similarity

Methods 이 논문이 사용한 방법론

ADOPT Please enter a description about the method here
Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

2024-03-05 · Philipp J. Rösch, Norbert Oswald, Michaela Geierhos, Jindřich Libovický

Current multimodal models leveraging contrastive learning often face limitations in developing fine-grained conceptual understanding. This is due to random negative samples during pretraining, causing almost exclusively …

Concept AlignmentContrastive LearningImage-text Retrieval

FG-CLIP: Fine-Grained Visual and Textual Alignment

2025-05-08 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…

Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions

2025-08-04 · Ziteng Wang, Siqi Yang, Limeng Qiao, Lin Ma arxiv

Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framew…

Fine-Grained Visual RecognitionContrastive LearningImage Manipulation

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce

2026-04-22 · Biao Zhang, Lixin Chen, Bin Zhang, Zongwei Wang 외 arxiv

Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. Large representation models (e.g., VLM2Vec) demonstrate strong multimodal understanding capabilities, yet they struggle with …

Representation LearningContrastive Learning

Improving fine-grained understanding in image-text pre-training

2024-01-18 · Ioana Bica, Anastasija Ilić, Matthias Bauer, Goker Erdogan 외

We introduce SPARse Fine-grained Contrastive Alignment (SPARC), a simple method for pretraining more fine-grained multimodal representations from image-text pairs. Given that multiple image patches often correspond to si…

object-detectionObject Detection