paper-with-me

Papers

SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment

2025-11-04 · Wenbo Lu arxiv

Vision-Language Pretraining (VLP) has achieved remarkable success across various downstream tasks, but such gains are largely driven by scaling up on training data. Yet, literature methods treat image-text pairs as isolated training examples; this neglects the rich relational structure naturally present in many domains, such as e-commerce product co-purchase graphs and social recommendation networks. Inspired by neuroscientific evidence that human encodes knowledge as relationship cognitive maps, we introduce Structure-aware Language-Image Pretraining (SLIP). SLIP integrates a structural contrastive loss to align modalities while also modeling relationships between neighboring entities in a structured graph. To support this paradigm, we construct a large-scale Amazon Product Co-purchase Multimodal Graph Dataset, enabling structured cross-modality supervision at scale. Experiment results show that SLIP consistently outperforms CLIP on cross-modal retrieval and classification tasks in both zero-shot and few-shot settings, showing the value of relational supervision for cross-modal alignment.

📄 PDF Abstract BibTeX arXiv:2511.03019

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal Retrieval

Similar Papers 제목 키워드 기반

Zero Shot Context-Based Object Segmentation using SLIP (SAM+CLIP)

2024-05-12 · Saaketh Koundinya Gundavarapu, Arushi Arora, Shreya Agarwal

We present SLIP (SAM+CLIP), an enhanced architecture for zero-shot object segmentation. SLIP combines the Segment Anything Model (SAM) \cite{kirillov2023segment} with the Contrastive Language-Image Pretraining (CLIP) \ci…

ObjectSegmentationSemantic Segmentation

SLIP: Spoof-Aware One-Class Face Anti-Spoofing with Language Image Pretraining

2025-03-25 · Pei-Kai Huang, Jun-Xiong Chong, Cheng-Hsuan Chiang, Tzu-Hsien Chen 외

Face anti-spoofing (FAS) plays a pivotal role in ensuring the security and reliability of face recognition systems. With advancements in vision-language pretrained (VLP) models, recent two-class FAS techniques have lever…

DisentanglementFace Anti-SpoofingFace Recognition

FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding

2025-11-18 · Zhenshi Li, Weikang Yu, Dilxat Muhtar, Xueliang Zhang 외 arxiv

As CLIP's global alignment limits its ability to capture fine-grained details, recent efforts have focused on enhancing its region-text alignment. However, current remote sensing (RS)-specific CLIP variants still inherit…

Semantic SegmentationText Retrieval

SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection

2026-05-22 · Chenxu Wang, Yuxuan Li, Yunheng Li, Xiang Li 외 arxiv

Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enumerating open-set categories via black-box data to acquire fine-graine…

Domain GeneralizationContrastive LearningObject Detection

Structure Learning of Probabilistic Logic Programs by Searching the Clause Space

2013-09-09 · Elena Bellodi, Fabrizio Riguzzi

Learning probabilistic logic programming languages is receiving an increasing attention and systems are available for learning the parameters (PRISM, LeProbLog, LFI-ProbLog and EMBLEM) or both the structure and the param…