paper-with-me

홈 › Papers

Task Adaptive Parameter Sharing for Multi-Task Learning

2022-03-30 · CVPR 2022 1 · Matthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran, Charless Fowlkes, Rahul Bhotika, Stefano Soatto

Adapting pre-trained models with broad capabilities has become standard practice for learning a wide range of downstream tasks. The typical approach of fine-tuning different models for each task is performant, but incurs a substantial memory cost. To efficiently learn multiple downstream tasks we introduce Task Adaptive Parameter Sharing (TAPS), a general method for tuning a base model to a new task by adaptively modifying a small, task-specific subset of layers. This enables multi-task learning while minimizing resources used and competition between tasks. TAPS solves a joint optimization problem which determines which layers to share with the base model and the value of the task-specific weights. Further, a sparsity penalty on the number of active layers encourages weight sharing with the base model. Compared to other methods, TAPS retains high accuracy on downstream tasks while introducing few task-specific parameters. Moreover, TAPS is agnostic to the model architecture and requires only minor changes to the training scheme. We evaluate our method on a suite of fine-tuning tasks and architectures (ResNet, DenseNet, ViT) and show that it achieves state-of-the-art performance while being simple to implement.

📄 PDF Abstract BibTeX arXiv:2203.16708

Code (1)

MattWallingford/TAPS pytorch

Tasks

Multi-Task Learning

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Batch Normalization 설명 없음
Dense Block A Dense Block is a module used in convolutional neural networks that connects *all layers* (with matching feature-map sizes) directly with each other. It was originally…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
BASE 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

ASLoRA: Adaptive Sharing Low-Rank Adaptation Across Layers

2024-12-13 · Junyan Hu, Xue Xiao, Mengqi Zhang, Yao Chen 외

As large language models (LLMs) grow in size, traditional full fine-tuning becomes increasingly impractical due to its high computational and storage costs. Although popular parameter-efficient fine-tuning methods, such …

parameter-efficient fine-tuning

Adaptive parameter sharing for multi-agent reinforcement learning

2023-12-14 · Dapeng Li, Na Lou, Bin Zhang, Zhiwei Xu 외

Parameter sharing, as an important technique in multi-agent systems, can effectively solve the scalability issue in large-scale agent problems. However, the effectiveness of parameter sharing largely depends on the envir…

DiversityMulti-agent Reinforcement Learningreinforcement-learningReinforcement Learning

Learning Sparse Sharing Architectures for Multiple Tasks

2019-11-12 · Tianxiang Sun, Yunfan Shao, Xiaonan Li, PengFei Liu 외

Most existing deep multi-task learning models are based on parameter sharing, such as hard sharing, hierarchical sharing, and soft sharing. How choosing a suitable sharing mechanism depends on the relations among the tas…

Multi-Task Learning

OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes

2023-08-23 · ICCV 2023 1 · Tao Xie, Kun Dai, Siyi Lu, Ke Wang 외

In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the t…

Multi-Task LearningVisual Localization

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection

2026-07-03 · Wen Dong, Zhao Wang, Shuangqing Zhang, Kai Sun 외 arxiv

Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowledge evolves. Existing continual learning methods often suffer from se…

parameter-efficient fine-tuningContinual LearningAnomaly Detection