paper-with-me

Papers

Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning

2026-03-06 · Haojie Pu, Zhuoming Li, Yongbiao Gao, Yuheng Jia arxiv

Generative zero-shot learning (ZSL) synthesizes features for unseen classes, leveraging semantic conditions to transfer knowledge from seen classes. However, it also introduces two intrinsic challenges: (1) class-level attributes fails to capture instance-specific visual appearances due to substantial intra-class variability, thus causing the class-instance gap; (2) the substantial mismatch between semantic and visual feature distributions, manifested in inter-class correlations, gives rise to the semantic-visual domain gap. To address these challenges, we propose an Attribute Distribution Modeling and Semantic-Visual Alignment (ADiVA) approach, jointly modeling attribute distributions and performing explicit semantic-visual alignment. Specifically, our ADiVA consists of two modules: an Attribute Distribution Modeling (ADM) module that learns a transferable attribute distribution for each class and samples instance-level attributes for unseen classes, and a Visual-Guided Alignment (VGA) module that refines semantic representations to better reflect visual structures. Experiments on three widely used benchmark datasets demonstrate that ADiVA significantly outperforms state-of-the-art methods (e.g., achieving gains of 4.7% and 6.1% on AWA2 and SUN, respectively). Moreover, our approach can serve as a plugin to enhance existing generative ZSL methods.

📄 PDF Abstract BibTeX arXiv:2603.06281

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-Shot Learning

Similar Papers 제목 키워드 기반

Segment-driven Structural Induction and Semantic Alignment for Heterogeneous Tabular Representation

2026-06-01 · Woojun Jung, Susik Yoon arxiv

Real-world domains often contain heterogeneous tables whose headers vary while their underlying attribute semantics are shared, making it difficult to induce domain-specialized semantics from table-local evidence alone. …

EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space

2026-01-17 · Jing Zhang, Bingjie Fan, Jixiang Zhu, Zhe Wang arxiv

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, a…

Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training

2025-08-01 · Weiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang 외 arxiv

Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a…

Contrastive LearningTransfer Learning

Object Attribute Matters in Visual Question Answering

2023-12-20 · Peize Li, Qingyi Si, Peng Fu, Zheng Lin 외

Visual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to…

AttributeGraph Neural NetworkKnowledge DistillationObject+5

CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning

2024-04-15 · Haojian Huang, Xiaozhen Qiao, Zhuo Chen, Haodong Chen 외

Zero-shot learning (ZSL) enables the recognition of novel classes by leveraging semantic knowledge transfer from known to unknown categories. This knowledge, typically encapsulated in attribute descriptions, aids in iden…

AttributeTransfer LearningVisual LocalizationZero-Shot Learning