paper-with-me

홈 › Papers

CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval

2025-10-07 · Bin Kang, Bin Chen, Junjie Wang, Yulin Li, Junzhi Zhao, Zhuotao Tian arxiv

Existing Visual Language Models (VLMs) suffer structural limitations where a few low contribution tokens may excessively capture global semantics, dominating the information aggregation process and suppressing the discriminative features in text-driven image retrieval tasks. To address this, we introduce \textbf{CalibCLIP}, a training-free method designed to calibrate the suppressive effect of dominant tokens. Specifically, in the visual space, we propose the Contrastive Visual Enhancer (CVE), which decouples visual features into target and low information regions. Subsequently, it identifies dominant tokens and dynamically suppresses their representations.In the textual space, we introduce the Discriminative Concept Calibrator (DCC), which aims to differentiate between general and discriminative concepts within the text query. By mitigating the challenges posed by generic concepts and improving the representations of discriminative concepts, DCC strengthens the differentiation among similar samples. Finally, extensive experiments demonstrate consistent improvements across seven benchmarks spanning three image retrieval tasks, underscoring the effectiveness of CalibCLIP. Code is available at: https://github.com/kangbin98/CalibCLIP

📄 PDF Abstract BibTeX arXiv:2510.05586

Code (0)

등록된 구현이 없습니다.

Tasks

Image Retrieval

Similar Papers 제목 키워드 기반

Calibration-Gated LLM Pseudo-Observations for Online Contextual Bandits

2026-04-16 · Maksim Pershin, Ivan Golovanov, Pavel Baltabaev, Natalia Trankova arxiv

Contextual bandit algorithms suffer from high regret during cold-start, when the learner has insufficient data to distinguish good arms from bad. We propose augmenting Disjoint LinUCB with LLM pseudo-observations: after …

A Cluster-based Approach for Improving Isotropy in Contextual Embedding Space

2021-06-02 · ACL 2021 5 · Sara Rajaee, Mohammad Taher Pilehvar

The representation degeneration problem in Contextual Word Representations (CWRs) hurts the expressiveness of the embedding space by forming an anisotropic cone where even unrelated words have excessively positive correl…

Open-Vocabulary Segmentation with Semantic-Assisted Calibration

2023-12-07 · CVPR 2024 1 · Yong liu, Sule Bai, Guanbin Li, Yitong Wang 외

This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding, alignment…

AttributeOpen Vocabulary Semantic Segmentation

Contextual Representation Learning beyond Masked Language Modeling

2022-04-08 · ACL 2022 5 · Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu, Hao Zhou 외

How do masked language models (MLMs) such as BERT learn contextual representations? In this work, we analyze the learning dynamics of MLMs. We find that MLMs adopt sampled embeddings as anchors to estimate and inject con…

Language ModelingLanguage ModellingMasked Language ModelingRepresentation Learning

Contextual Representation Learning beyond Masked Language Modeling

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Currently, masked language modeling (e.g., BERT) is the prime choice to learn contextualized representations. Due to the pervasiveness, it naturally raises an interesting question: how do masked language models (MLMs) le…

Language ModelingLanguage ModellingMasked Language ModelingRepresentation Learning