paper-with-me

홈 › Papers

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

2026-01-13 · Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin arxiv

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines.

📄 PDF Abstract BibTeX arXiv:2601.08476

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-modal Dynamic Proxy Learning for Personalized Multiple Clustering

2025-11-10 · Jinfeng Xu, Zheyu Chen, Shuo Yang, Jinze Li 외 arxiv

Multiple clustering aims to discover diverse latent structures from different perspectives, yet existing methods generate exhaustive clusterings without discerning user interest, necessitating laborious manual screening.…

CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising

2021-12-14 · Jianjie Luo, Yehao Li, Yingwei Pan, Ting Yao 외

BERT-type structure has led to the revolution of vision-language pre-training and the achievement of state-of-the-art results on numerous vision-language downstream tasks. Existing solutions dominantly capitalize on the …

Cross-Modal RetrievalDecoderDenoisingLanguage Modeling+6

Intra-Modal Proxy Learning for Zero-Shot Visual Categorization with CLIP

2023-09-21 · NeurIPS 2023 11

Vision-language pre-training methods, e.g., CLIP, demonstrate an impressive zero-shot performance on visual categorizations with the class proxy from the text embedding of the class name. However, the modality gap betwee…

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

2026-02-11 · Jinjie Shen, Jing Wu, Yaxiong Wang, Lechao Cheng 외 arxiv

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, this paper ta…

Reinforcement Learning

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning

2026-06-30 · Hongyi Lin, Yang Liu, Jinhua Zhao, Xiaobo Qu arxiv

Foundation models are increasingly integrated into embodied intelligence systems, but directly assigning them structured prediction tasks requires precise geometric and numerical estimation, where specialized models ofte…

Structured PredictionTrajectory PredictionSemantic Segmentation2D Object Detection