paper-with-me

홈 › Papers

Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition

2024-07-19 · Yurong Zhang, Honghao Chen, Xinyu Zhang, Xiangxiang Chu, Li Song

Parameter-efficient transfer learning (PETL) is a promising task, aiming to adapt the large-scale pre-trained model to downstream tasks with a relatively modest cost. However, current PETL methods struggle in compressing computational complexity and bear a heavy inference burden due to the complete forward process. This paper presents an efficient visual recognition paradigm, called Dynamic Adapter (Dyn-Adapter), that boosts PETL efficiency by subtly disentangling features in multiple levels. Our approach is simple: first, we devise a dynamic architecture with balanced early heads for multi-level feature extraction, along with adaptive training strategy. Second, we introduce a bidirectional sparsity strategy driven by the pursuit of powerful generalization ability. These qualities enable us to fine-tune efficiently and effectively: we reduce FLOPs during inference by 50%, while maintaining or even yielding higher recognition accuracy. Extensive experiments on diverse datasets and pretrained backbones demonstrate the potential of Dyn-Adapter serving as a general efficiency booster for PETL in vision recognition tasks.

📄 PDF Abstract BibTeX arXiv:2407.14302

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

D$^2$ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition

2023-12-03 · Wenjie Pei, Qizhong Tan, Guangming Lu, Jiandong Tian

Adapting large pre-trained image models to few-shot action recognition has proven to be an effective and efficient strategy for learning robust feature extractors, which is essential for few-shot learning. Typical fine-t…

Action RecognitionFew-Shot action recognitionFew Shot Action RecognitionFew-Shot Learning

Disentangling Semantic-to-visual Confusion for Zero-shot Learning

2021-06-16 · Zihan Ye, Fuyuan Hu, Fan Lyu, Linyan Li 외

Using generative models to synthesize visual features from semantic distribution is one of the most popular solutions to ZSL image classification in recent years. The triplet loss (TL) is popularly used to generate reali…

Generative Adversarial Networkimage-classificationImage ClassificationTriplet+1

Preventing Shortcuts in Adapter Training via Providing the Shortcuts

2025-10-23 · Anujraaj Argo Goyal, Guocheng Gordon Qian, Huseyin Coskun, Aarush Gupta 외 arxiv

Adapter-based training has emerged as a key mechanism for extending the capabilities of powerful foundation image generators, enabling personalized and stylized text-to-image synthesis. These adapters are typically train…

Image Reconstruction

Cross-composition Feature Disentanglement for Compositional Zero-shot Learning

2024-08-19 · Yuxia Geng, Runkai Zhu, Jiaoyan Chen, Jintai Chen 외

Disentanglement of visual features of primitives (i.e., attributes and objects) has shown exceptional results in Compositional Zero-shot Learning (CZSL). However, due to the feature divergence of an attribute (resp. obje…

AttributeCompositional Zero-Shot LearningDisentanglementLanguage Modeling+2

A Deeper Look at the Unsupervised Learning of Disentangled Representations in $β$-VAE from the Perspective of Core Object Recognition

2020-04-25 · Harshvardhan Sikka

The ability to recognize objects despite there being differences in appearance, known as Core Object Recognition, forms a critical part of human perception. While it is understood that the brain accomplishes Core Object …

Bayesian InferenceObjectObject RecognitionVariational Inference