paper-with-me

홈 › Papers

Towards a Unified View of Parameter-Efficient Transfer Learning

2021-10-08 · ICLR 2022 4 · Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, Graham Neubig

Fine-tuning large pre-trained language models on downstream tasks has become the de-facto learning paradigm in NLP. However, conventional approaches fine-tune all the parameters of the pre-trained model, which becomes prohibitive as the model size and the number of tasks grow. Recent work has proposed a variety of parameter-efficient transfer learning methods that only fine-tune a small number of (extra) parameters to attain strong performance. While effective, the critical ingredients for success and the connections among the various methods are poorly understood. In this paper, we break down the design of state-of-the-art parameter-efficient transfer learning methods and present a unified framework that establishes connections between them. Specifically, we re-frame them as modifications to specific hidden states in pre-trained models, and define a set of design dimensions along which different methods vary, such as the function to compute the modification and the position to apply the modification. Through comprehensive empirical studies across machine translation, text summarization, language understanding, and text classification benchmarks, we utilize the unified view to identify important design choices in previous methods. Furthermore, our unified framework enables the transfer of design elements across different approaches, and as a result we are able to instantiate new parameter-efficient fine-tuning methods that tune less parameters than previous methods while being more effective, achieving comparable results to fine-tuning all parameters on all four tasks.

📄 PDF Abstract BibTeX arXiv:2110.04366

Code (1)

jxhe/unify-parameter-efficient-tuning 공식 구현 jax

Tasks

Machine Translationparameter-efficient fine-tuningtext-classificationText ClassificationText SummarizationTransfer Learning

Similar Papers 제목 키워드 기반

A PAC-Bayesian bound for Lifelong Learning

2013-11-12 · Anastasia Pentina, Christoph H. Lampert

Transfer learning has received a lot of attention in the machine learning community over the last years, and several effective algorithms have been developed. However, relatively little is known about their theoretical p…

Lifelong learningTransfer Learning

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model

2026-07-13 · Xinghang Li, Jun Guo, Qiwei Li, Long Qian 외 hf

Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coh…

Text-to-Image GenerationScene GenerationVideo GenerationImage Editing

Towards a Unified View on Visual Parameter-Efficient Transfer Learning

2022-10-03 · Bruce X. B. Yu, Jianlong Chang, Lingbo Liu, Qi Tian 외

Parameter efficient transfer learning (PETL) aims at making good use of the representation knowledge in the pre-trained large models by fine-tuning a small number of parameters. Recently, taking inspiration from the natu…

Action RecognitionImage ClassificationTransfer LearningVideo Recognition

UPGPT: Universal Diffusion Model for Person Image Generation, Editing and Pose Transfer

2023-04-18 · Soon Yau Cheong, Armin Mustafa, Andrew Gilbert

Text-to-image models (T2I) such as StableDiffusion have been used to generate high quality images of people. However, due to the random nature of the generation process, the person has a different appearance e.g. pose, f…

DisentanglementImage GenerationPose TransferText-to-Image Generation+1

Parameter-Efficient Fine-Tuning for Pre-Trained Vision Models: A Survey

2024-02-03 · Yi Xin, Jianjiang Yang, Siqi Luo, Haodi Zhou 외

Large-scale pre-trained vision models (PVMs) have shown great potential for adaptability across various downstream vision tasks. However, with state-of-the-art PVMs growing to billions or even trillions of parameters, th…

parameter-efficient fine-tuningTransfer Learning