paper-with-me

Papers

PSVMA+: Exploring Multi-granularity Semantic-visual Adaption for Generalized Zero-shot Learning

2024-10-15 · Man Liu, Huihui Bai, Feng Li, Chunjie Zhang, Yunchao Wei, Meng Wang, Tat-Seng Chua, Yao Zhao

Generalized zero-shot learning (GZSL) endeavors to identify the unseen categories using knowledge from the seen domain, necessitating the intrinsic interactions between the visual features and attribute semantic features. However, GZSL suffers from insufficient visual-semantic correspondences due to the attribute diversity and instance diversity. Attribute diversity refers to varying semantic granularity in attribute descriptions, ranging from low-level (specific, directly observable) to high-level (abstract, highly generic) characteristics. This diversity challenges the collection of adequate visual cues for attributes under a uni-granularity. Additionally, diverse visual instances corresponding to the same sharing attributes introduce semantic ambiguity, leading to vague visual patterns. To tackle these problems, we propose a multi-granularity progressive semantic-visual mutual adaption (PSVMA+) network, where sufficient visual elements across granularity levels can be gathered to remedy the granularity inconsistency. PSVMA+ explores semantic-visual interactions at different granularity levels, enabling awareness of multi-granularity in both visual and semantic elements. At each granularity level, the dual semantic-visual transformer module (DSVTM) recasts the sharing attributes into instance-centric attributes and aggregates the semantic-related visual regions, thereby learning unambiguous visual features to accommodate various instances. Given the diverse contributions of different granularities, PSVMA+ employs selective cross-granularity learning to leverage knowledge from reliable granularities and adaptively fuses multi-granularity features for comprehensive representations. Experimental results demonstrate that PSVMA+ consistently outperforms state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2410.11560

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeDiversityGeneralized Zero-Shot LearningZero-Shot Learning

Similar Papers 제목 키워드 기반

Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot Learning

2023-03-27 · CVPR 2023 1 · Man Liu, Feng Li, Chunjie Zhang, Yunchao Wei 외

Generalized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize…

AttributeDecoderGeneralized Zero-Shot LearningZero-Shot Learning

Granulon: Awakening Pixel-Level Visual Encoders with Adaptive Multi-Granularity Semantics for MLLM

2026-03-09 · Junyuan Mao, Qiankun Li, Linghao Meng, Zhicheng He 외 arxiv

Recent advances in multimodal large language models largely rely on CLIP-based visual encoders, which emphasize global semantic alignment but struggle with fine-grained visual understanding. In contrast, DINOv3 provides …

Multi-Granularity Mutual Refinement Network for Zero-Shot Learning

2025-11-11 · Ning Wang, Long Yu, Cong Hua, Guangming Zhu 외 arxiv

Zero-shot learning (ZSL) aims to recognize unseen classes with zero samples by transferring semantic knowledge from seen classes. Current approaches typically correlate global visual features with semantic information (i…

Zero-Shot Learning

An Information Compensation Framework for Zero-Shot Skeleton-based Action Recognition

2024-06-02 · Haojun Xu, Yan Gao, Jie Li, Xinbo Gao

Zero-shot human skeleton-based action recognition aims to construct a model that can recognize actions outside the categories seen during training. Previous research has focused on aligning sequences' visual and semantic…

Action RecognitionEnsemble LearningSkeleton Based Action RecognitionZero-Shot Action Recognition+1

Semantic-SAM: Segment and Recognize Anything at Any Granularity

2023-07-10 · Feng Li, Hao Zhang, Peize Sun, Xueyan Zou 외

In this paper, we introduce Semantic-SAM, a universal image segmentation model to enable segment and recognize anything at any desired granularity. Our model offers two key advantages: semantic-awareness and granularity-…

Image SegmentationSegmentationSemantic Segmentation