paper-with-me

홈 › Papers

Cross-composition Feature Disentanglement for Compositional Zero-shot Learning

2024-08-19 · Yuxia Geng, Runkai Zhu, Jiaoyan Chen, Jintai Chen, Zhuo Chen, Xiang Chen, Can Xu, Yuxiang Wang, Xiaoliang Xu

Disentanglement of visual features of primitives (i.e., attributes and objects) has shown exceptional results in Compositional Zero-shot Learning (CZSL). However, due to the feature divergence of an attribute (resp. object) when combined with different objects (resp. attributes), it is challenging to learn disentangled primitive features that are general across different compositions. To this end, we propose the solution of cross-composition feature disentanglement, which takes multiple primitive-sharing compositions as inputs and constrains the disentangled primitive features to be general across these compositions. More specifically, we leverage a compositional graph to define the overall primitive-sharing relationships between compositions, and build a task-specific architecture upon the recently successful large pre-trained vision-language model (VLM) CLIP, with dual cross-composition disentangling adapters (called L-Adapter and V-Adapter) inserted into CLIP's frozen text and image encoders, respectively. Evaluation on three popular CZSL benchmarks shows that our proposed solution significantly improves the performance of CZSL, and its components have been verified by solid ablation studies.

📄 PDF Abstract BibTeX arXiv:2408.09786

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeCompositional Zero-Shot LearningDisentanglementLanguage ModelingLanguage ModellingZero-Shot Learning

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Compositional Zero-Shot Learning: A Survey

2025-10-13 · Ans Munir, Faisal Z. Qureshi, Mohsen Ali, Muhammad Haris Khan arxiv

Compositional Zero-Shot Learning (CZSL) is a critical task in computer vision that enables models to recognize unseen combinations of known attributes and objects during inference, addressing the combinatorial challenge …

Compositional Zero-Shot Learning

CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement

2025-11-20 · Pan Yang, Cheng Deng, Jing Yang, Han Zhao 외 arxiv

Compositional zero-shot learning (CZSL) aims to learn the concepts of attributes and objects in seen compositions and to recognize their unseen compositions. Most Contrastive Language-Image Pre-training (CLIP)-based CZSL…

Compositional Zero-Shot Learning

Learning Attention as Disentangler for Compositional Zero-shot Learning

2023-03-27 · CVPR 2023 1 · Shaozhe Hao, Kai Han, Kwan-Yee K. Wong

Compositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the dis…

AttributeCompositional Zero-Shot LearningDisentanglementZero-Shot Learning

Recursive Disentanglement Network

2021-09-29 · ICLR 2022 4 · Yixuan Chen, Yubin Shi, Dongsheng Li, Yujiang Wang 외

Disentangled feature representation is essential for data-efficient learning. The feature space of deep models is inherently compositional. Existing $\beta$-VAE-based methods, which only apply disentanglement regularizat…

DisentanglementInductive BiasRepresentation Learning

Visual Referential Games Further the Emergence of Disentangled Representations

2023-04-27 · Kevin Denamganaï, Sondess Missaoui, James Alfred Walker

Natural languages are powerful tools wielded by human beings to communicate information. Among their desirable properties, compositionality has been the main focus in the context of referential games and variants, as it …

DisentanglementInformativeness