paper-with-me

Papers

HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities

2026-05-06 · Esra Dönmez, Pascal Tilli, Hsiu-Yu Yang, Thang Vu, Carina Silberer arxiv

Image-Text-Matching (ITM) is one of the defacto methods of learning generalized representations from a large corpus in Vision and Language (VL). However, due to the weak association between the web-collected image-text pairs, models fail to show a fine-grained understanding of the combined semantics of these modalities. To address this issue we propose Hard Negative Captions (HNC): an automatically created dataset containing foiled hard negative captions for ITM training towards achieving fine-grained cross-modal comprehension in VL. Additionally, we provide a challenging manually-created test set for benchmarking models on a fine-grained cross-modal mismatch task with varying levels of compositional complexity. Our results show the effectiveness of training on HNC by improving the models' zero-shot capabilities in detecting mismatches on diagnostic tasks and performing robustly under noisy visual input scenarios. Also, we demonstrate that HNC models yield a comparable or better initialization for fine-tuning

📄 PDF Abstract BibTeX arXiv:2605.06157

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Enhancing Conceptual Understanding in Multimodal Contrastive Learning through Hard Negative Samples

2024-03-05 · Philipp J. Rösch, Norbert Oswald, Michaela Geierhos, Jindřich Libovický

Current multimodal models leveraging contrastive learning often face limitations in developing fine-grained conceptual understanding. This is due to random negative samples during pretraining, causing almost exclusively …

Concept AlignmentContrastive LearningImage-text Retrieval

FG-CLIP: Fine-Grained Visual and Textual Alignment

2025-05-08 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short c…

Image-text Retrievalobject-detectionObject DetectionOpen-vocabulary object detection+5

SPOT! Revisiting Video-Language Models for Event Understanding

2023-11-21 · Gengyuan Zhang, Jinhe Bi, Jindong Gu, Yanyu Chen 외

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint repre…

AttributeVideo Understanding

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions

2025-08-04 · Ziteng Wang, Siqi Yang, Limeng Qiao, Lin Ma arxiv

Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framew…

Fine-Grained Visual RecognitionContrastive LearningImage Manipulation

Can Text-to-Video Generation help Video-Language Alignment?

2025-03-24 · CVPR 2025 1 · Luca Zanella, Massimiliano Mancini, Willi Menapace, Sergey Tulyakov 외

Recent video-language alignment models are trained on sets of videos, each with an associated positive caption and a negative caption generated by large language models. A problem with this procedure is that negative cap…

Text-to-Video GenerationVideo Generation