paper-with-me

홈 › Papers

Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image Pre-training

2023-01-01 · ICCV 2023 1 · Bumsoo Kim, Yeonsik Jo, Jinhyung Kim, Seunghwan Kim

Contrastive Language-Image Pretraining has emerged as a prominent approach for training vision and text encoders with uncurated image-text pairs from the web. To enhance data-efficiency, recent efforts have introduced additional supervision terms that involve random-augmented views of the image. However, since the image augmentation process is unaware of its text counterpart, this procedure could cause various degrees of image-text misalignments during training. Prior methods either disregarded this discrepancy or introduced external models to mitigate the impact of misalignments during training. In contrast, we propose a novel metric learning approach that capitalizes on these misalignments as an additional training source, which we term "Misalign, Contrast then Distill (MCD)". Unlike previous methods that treat augmented images and their text counterparts as simple positive pairs, MCD predicts the continuous scales of misalignment caused by the augmentation. Our extensive experimental results show that our proposed MCD achieves state-of-the-art transferability in multiple classification and retrieval downstream datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image AugmentationMetric LearningRetrieval

Similar Papers 제목 키워드 기반

Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image Pretraining

2023-12-19 · Bumsoo Kim, Yeonsik Jo, Jinhyung Kim, Seung Hwan Kim

Contrastive Language-Image Pretraining has emerged as a prominent approach for training vision and text encoders with uncurated image-text pairs from the web. To enhance data-efficiency, recent efforts have introduced ad…

Image AugmentationMetric LearningRetrieval

Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning

2024-05-23 · Hector Kohler, Quentin Delfosse, Riad Akrour, Kristian Kersting 외

Deep reinforcement learning agents are prone to goal misalignments. The black-box nature of their policies hinders the detection and correction of such misalignments, and the trust necessary for real-world deployment. So…

Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning

Addressing and Visualizing Misalignments in Human Task-Solving Trajectories

2024-09-21 · Sejin Kim, Hosung Lee, Sundong Kim

Understanding misalignments in human task-solving trajectories is crucial for enhancing AI models trained to closely mimic human reasoning. This study categorizes such misalignments into three types: (1) lack of function…

Misalignment Resilient Diffractive Optical Networks

2020-05-23 · Deniz Mengu, Yifan Zhao, Nezih T. Yardimci, Yair Rivenson 외

As an optical machine learning framework, Diffractive Deep Neural Networks (D2NN) take advantage of data-driven training methods used in deep learning to devise light-matter interaction in 3D for performing a desired sta…

Object Recognition

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

2026-07-16 · Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier, Akshay Chaudhari 외 arxiv

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic mi…