paper-with-me

홈 › Papers

TIER: Text-Image Entropy Regularization for CLIP-style models

2022-12-13 · Anil Palepu, Andrew L. Beam

In this paper, we introduce a novel regularization scheme on contrastive language-image pre-trained (CLIP) medical vision models. Our approach is based on the observation that on many medical imaging tasks text tokens should only describe a small number of image regions and, likewise, each image region should correspond to only a few text tokens. In CLIP-style models, this implies that text-token embeddings should have high similarity to only a small number of image-patch embeddings for a given image-text pair. We formalize this observation using a novel regularization scheme that penalizes the entropy of the text-token to image-patch similarity scores. We qualitatively and quantitatively demonstrate that the proposed regularization scheme shrinks most of the pairwise text-token and image-patch similarity scores towards zero, thus achieving the desired effect. We demonstrate the promise of our approach in an important medical context, chest x-rays, where this underlying sparsity hypothesis naturally arises. Using our proposed approach, we achieve state of the art (SOTA) average zero-shot performance on the CheXpert and Padchest chest x-ray datasets, outperforming an unregularized version of the model and several recently published self-supervised models.

📄 PDF Abstract BibTeX arXiv:2212.06710

Code (1)

apalepu13/tier_regularized_clip 공식 구현 pytorch

Similar Papers 제목 키워드 기반

TIER-A: Denoising Learning Framework for Information Extraction

2022-11-13 · Yongkang Li, Ming Zhang

With the development of deep neural language models, great progress has been made in information extraction recently. However, deep learning models often overfit on noisy data points, leading to poor performance. In this…

Denoising

Class Interference Regularization

2020-09-04 · Bharti Munjal, Sikandar Amin, Fabio Galasso

Contrastive losses yield state-of-the-art performance for person re-identification, face verification and few shot learning. They have recently outperformed the cross-entropy loss on classification at the ImageNet scale …

Face VerificationFew-Shot LearningPerson Re-IdentificationPerson Search+1

A Framework for Benchmarking Fairness-Utility Trade-offs in Text-to-Image Models via Pareto Frontiers

2025-08-22 · Marco N. Bochernitsan, Rodrigo C. Barros, Lucas S. Kupssinskü arxiv

Achieving fairness in text-to-image generation demands mitigating social biases without compromising visual fidelity, a challenge critical to responsible AI. Current fairness evaluation procedures for text-to-image model…

Text-to-Image Generation

AspectCLIP: Optimizing CLIP Representation Space via Aspect-Guided Consistency Regularization

2026-07-15 · Yiyang Yao, Shanglin Liu, Jianming Lv, Chengjun Wang 외 arxiv

Contrastive Language-Image Pretraining learns a shared representation space through large-scale contrastive learning. However, existing methods that enforce global consistency regularization overlook a key challenge: the…

Contrastive Learning

Dissecting CLIP: Decomposition with a Schur Complement-based Approach

2024-12-24 · Azim Ospanov, Mohammad Jalali, Farzan Farnia

The use of CLIP embeddings to assess the alignment of samples produced by text-to-image generative models has been extensively explored in the literature. While the widely adopted CLIPScore, derived from the cosine simil…

Diversity