paper-with-me

Papers

KODA: Contrastive Representation Comparison and Alignment for Vision-Language Foundation Models

2026-06-02 · Youqi Wu, Mohammad Jalali, Farzan Farnia arxiv

Vision-language foundation models such as CLIP and SigLIP provide widely used representations for multimodal learning systems. While these models are typically compared through downstream performance, such evaluations often do not explain how their representations differ structurally. In this work, we study this problem through the task of Contrastive Embedding Clustering: identifying sample subsets that are weakly clustered under one representation but strongly clustered under another. We propose \emph{Kernel Optimization for Discrepancy Analysis (KODA)}, a kernel-based framework for contrastive representation comparison and alignment. KODA constructs unified multimodal kernels through modality-wise kernel composition and formulates discrepancy discovery as a constrained optimization problem that searches for coherent structures in one representation while suppressing coherence in a reference representation. This yields interpretable discrepancy directions associated with specific sample subsets and modality interactions. To scale KODA to large vision-language datasets, we develop randomized low-dimensional approximations of joint kernels using random projections, including Random Fourier Features for shift-invariant kernels. Empirically, KODA identifies consistent and interpretable discrepancy structures across vision-language representations and provides sample subsets for representation alignment. The code is available at https://github.com/yokiwuuu/KODA.

📄 PDF Abstract BibTeX arXiv:2606.04180

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Barycentric alignment for instance-level comparison of neural representations

2026-02-09 · Shreya Saha, Zoe Wanying He, Meenakshi Khosla arxiv

Comparing representations across neural networks is challenging because representations admit symmetries, such as arbitrary reordering of units or rotations of activation space, that obscure underlying equivalence betwee…

Time Series, Vision, and Language: Exploring the Limits of Alignment in Contrastive Representation Spaces

2026-02-22 · Pratham Yashwante, Rose Yu arxiv

The Platonic Representation Hypothesis posits that learned representations from models trained on different modalities converge to a shared latent structure of the world. However, this hypothesis has largely been examine…

Contrastive Learning

Recent advances on inconsistency indices for pairwise comparisons - a commentary

2015-03-28 · Matteo Brunelli

This paper recalls the definition of consistency for pairwise comparison matrices and briefly presents the concept of inconsistency index in connection to other aspects of the theory of pairwise comparisons. By commentin…

Vision-Language Pre-Training with Triple Contrastive Learning

2022-02-21 · CVPR 2022 1 · Jinyu Yang, Jiali Duan, Son Tran, Yi Xu 외

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the…

Contrastive Learningcross-modal alignmentCross-Modal RetrievalImage-text Retrieval+7

Spatio-temporal Contrastive Domain Adaptation for Action Recognition

2021-06-19 · CVPR 2021 1 · Xiaolin Song, Sicheng Zhao, Jingyu Yang, Huanjing Yue 외

Unsupervised domain adaptation (UDA) for human action recognition is a practical and challenging problem. Compared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial rep…

Action RecognitionContrastive LearningDomain AdaptationSelf-Supervised Learning+2