"Does it come in black?" CLIP-like models are zero-shot recommenders
Product discovery is a crucial component for online shopping. However, item-to-item recommendations today do not allow users to explore changes along selected dimensions: given a query item, can a model suggest something similar but in a different color? We consider item recommendations of the comparative nature (e.g. "something darker") and show how CLIP-based models can support this use case in a zero-shot manner. Leveraging a large model built for fashion, we introduce GradREC and its industry potential, and offer a first rounded assessment of its strength and weaknesses.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
“Does it come in black?” CLIP-like models are zero-shot recommenders
Product discovery is a crucial component for online shopping. However, item-to-item recommendations today do not allow users to explore changes along selected dimensions: given a query item, can a model suggest something…
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
As object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) dete…
ObjectOut of Distribution (OOD) DetectionCAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning
In this paper, we study the problem of Compositional Zero-Shot Learning (CZSL), which is to recognize novel attribute-object combinations with pre-existing concepts. Recent researchers focus on applying large-scale Visio…
AttributeCompositional Zero-Shot LearningZero-Shot LearningCLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
In this paper, we aim to build an adversarially robust zero-shot image classifier. We ground our work on CLIP, a vision-language pre-trained encoder model that can perform zero-shot classification by matching an image wi…
Denoisingzero-shot-classificationZero-Shot LearningDoes CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
Foundation models like CLIP are trained on hundreds of millions of samples and effortlessly generalize to new tasks and inputs. Out of the box, CLIP shows stellar zero-shot and few-shot capabilities on a wide range of ou…
AttributeOut-of-Distribution Generalization