paper-with-me

Papers

Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

2024-03-05 · Philipp J. Rösch, Norbert Oswald, Michaela Geierhos, Jindřich Libovický

Current multimodal models leveraging contrastive learning often face limitations in developing fine-grained conceptual understanding. This is due to random negative samples during pretraining, causing almost exclusively very dissimilar concepts to be compared in the loss function. Consequently, the models struggle with fine-grained semantic differences. To address this problem, we introduce a novel pretraining method incorporating synthetic hard negative text examples. The hard negatives permute terms corresponding to visual concepts, leading to a more fine-grained visual and textual concept alignment. Further, we introduce InpaintCOCO, a new challenging dataset for assessing the fine-grained alignment of colors, objects, and sizes in vision-language models. We created the dataset using generative inpainting from COCO images by changing the visual concepts so that the images no longer match their original captions. Our results show significant improvements in fine-grained concept understanding across a wide range of vision-language datasets, including our InpaintCOCO dataset.

📄 PDF Abstract BibTeX arXiv:2403.02875

Code (0)

등록된 구현이 없습니다.

Tasks

Concept AlignmentContrastive LearningImage-text Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities

2022-11-12 · Zhongzhi Chen, Guang Liu, Bo-Wen Zhang, Fulong Ye 외

In this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model. Starting from the pre-trained multimodal representation model CLIP released by…

Contrastive LearningCross-Modal RetrievalImage ClassificationImage Retrieval+9

Multimodal contrastive learning for spatial gene expression prediction using histology images

2024-07-11 · Wenwen Min, Zhiceng Shi, Jun Zhang, Jun Wan 외

In recent years, the advent of spatial transcriptomics (ST) technology has unlocked unprecedented opportunities for delving into the complexities of gene expression patterns within intricate biological systems. Despite i…

Contrastive Learningwhole slide images

Towards Interpretable Hallucination Analysis and Mitigation in LVLMs via Contrastive Neuron Steering

2026-01-31 · Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu 외 arxiv

LVLMs achieve remarkable multimodal understanding and generation but remain susceptible to hallucinations. Existing mitigation methods predominantly focus on output-level adjustments, leaving the internal mechanisms that…

Visual Grounding

Enhancing Multimodal Compositional Reasoning of Visual Language Models with Generative Negative Mining

2023-11-07 · Ugur Sahin, Hang Li, Qadeer Khan, Daniel Cremers 외

Contemporary large-scale visual language models (VLMs) exhibit strong representation capacities, making them ubiquitous for enhancing image and text understanding tasks. They are often trained in a contrastive manner on …

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models

2026-01-30 · Yuansheng Gao, Jinman Zhao, Tong Zhang, Xingguo Xu 외 arxiv

Although Video Large Multimodal Models have achieved strong performance in video understanding, they still suffer from hallucination. Existing inference-time intervention methods usually modify videos under the contrasti…