paper-with-me

홈 › Papers

DINORANKCLIP: DINOv3 Distillation and Injection for Vision-Language Pretraining with High-Order Ranking Consistency

2026-05-07 · Shuyang Jiang, Nan Yu, Yiming Zhang, Zenghui Ding, Zhenyu Wu arxiv

Contrastive language-image pretraining (CLIP) suffers from two structural weaknesses: the symmetric InfoNCE loss discards the relative ordering among unmatched in-batch pairs, and global pooling collapses the visual representation into a semantic bottleneck that is poorly sensitive to fine-grained local structure. RANKCLIP partially addresses the first issue with a list-wise Plackett-Luce ranking-consistency loss, but its model is strictly first-order and inherits the second weakness untouched. We propose DINORANKCLIP, a pretraining framework that addresses both jointly. Our principal contribution is injecting a frozen DINOv3 teacher into the contrastive trunk through a dual-branch lightweight student and a multi-scale fusion module with channel-spatial attention, a self-attention refiner, and a conflict-aware gate that preserves the cross-modal alignment up to first order. Complementarily, we introduce a high-order Plackett-Luce ranking model in which the per-position utility is augmented with attention-parameterised pairwise and tuple-wise transition terms; the family contains CLIP and RANKCLIP as nested zero-order and first-order special cases, and the optimal order on every benchmark is $R^*=3$. The full empirical study -- order sweep, Fine-grained Probe on five datasets, four-node Modality-Gap analysis, six-variant Fusion ablation -- fits in 72 hours on a single eight-GPU H100 node and trains entirely on Conceptual Captions 3M. DINORANKCLIP consistently outperforms CLIP, CyCLIP, ALIP, and RANKCLIP under matched compute, with the largest relative gains on the fine-grained and out-of-distribution evaluations that most directly stress local structural reasoning.

📄 PDF Abstract BibTeX arXiv:2605.06592

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Semantically Compatible Knowledge Distillation for Cross-Domain Object Detection with Vision Foundation Models

2026-08-21 · Qifeng Zhang, Ting Xiang, Zeyuan Bai, Changjian Chen arxiv

Vision foundation models (VFMs) offer strong generalization capabilities for domain-adaptive object detection (DAOD). However, existing VFM-based methods overlook the spatial-scale discrepancy between teacher and student…

Knowledge DistillationObject Detection

X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning

2026-01-16 · Maanping Shao, Feihong Zhang, Gu Zhang, Baiye Cheng 외 arxiv

Visuomotor policies often leverage large pre-trained Vision Transformers (ViTs) for their powerful generalization capabilities. However, their significant data requirements present a major challenge in the data-scarce co…

Knowledge Distillation

Leveraging Foundation Models via Knowledge Distillation in Multi-Object Tracking: Distilling DINOv2 Features to FairMOT

2024-07-25 · Niels G. Faber, Seyed Sahand Mohammadi Ziabari, Fatemeh Karimi Nejadasl

Multiple Object Tracking (MOT) is a computer vision task that has been employed in a variety of sectors. Some common limitations in MOT are varying object appearances, occlusions, or crowded scenes. To address these chal…

Knowledge DistillationMulti-Object TrackingMultiple Object TrackingObject Tracking

FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision-Language Models

2026-01-29 · Chenyu Huang, Peng Ye, Xudong Tan, Jinhan Mu 외 arxiv

Efficiently enhancing the reasoning capabilities of Vision-Language Models (VLMs) by merging them with Large Reasoning Models (LRMs) has emerged as a promising direction. However, existing methods typically operate at a …

Accessing Vision Foundation Models at ImageNet-level Costs

2024-07-15 · Yitian Zhang, Xu Ma, Yue Bai, Huan Wang 외

Vision foundation models are renowned for their generalization ability due to massive training data. Nevertheless, they demand tremendous training resources, and the training data is often inaccessible, e.g., CLIP, DINOv…

Knowledge DistillationTransfer Learning