paper-with-me

홈 › Papers

Task-Attentive Transformer Architecture for Continual Learning of Vision-and-Language Tasks Using Knowledge Distillation

2023-03-25 · Yuliang Cai, Jesse Thomason, Mohammad Rostami

The size and the computational load of fine-tuning large-scale pre-trained neural network are becoming two major obstacles in adopting machine learning in many applications. Continual learning (CL) can serve as a remedy through enabling knowledge-transfer across sequentially arriving tasks which relaxes the need to fine-tune all network weights from scratch. However, existing CL algorithms primarily consider learning unimodal vision-only or language-only tasks. We develop a transformer-based CL architecture for learning bimodal vision-and-language tasks based on increasing the number of the learnable parameters dynamically and using knowledge distillation. The new additional parameters are used to specialize the network for each task. Our approach enables sharing information between the tasks while addressing the challenge of catastrophic forgetting. Our approach is scalable learning to a large number of tasks because it requires little memory and time overhead. Our model reaches state-of-the-art performance on challenging vision-and-language tasks.

📄 PDF Abstract BibTeX arXiv:2303.14423

Code (0)

등록된 구현이 없습니다.

Tasks

Continual LearningKnowledge DistillationTransfer Learning

Similar Papers 제목 키워드 기반

Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks

2024-01-27 · Yuliang Cai, Mohammad Rostami

Transformer neural networks are increasingly replacing prior architectures in a wide range of applications in different data modalities. The increasing size and computational demands of fine-tuning large pre-trained tran…

Continual LearningEdge-computingKnowledge Distillation

CvT-ASSD: Convolutional vision-Transformer Based Attentive Single Shot MultiBox Detector

2021-10-24 · Weiqiang Jin, Hang Yu

Due to the success of Bidirectional Encoder Representations from Transformers (BERT) in natural language process (NLP), the multi-head attention transformer has been more and more prevalent in computer-vision researches …

Computational EfficiencyNovel Object Detectionobject-detectionObject Detection+1

Continual Attentive Fusion for Incremental Learning in Semantic Segmentation

2022-02-01 · Guanglei Yang, Enrico Fini, Dan Xu, Paolo Rota 외

Over the past years, semantic segmentation, as many other tasks in computer vision, benefited from the progress in deep neural networks, resulting in significantly improved performance. However, deep architectures traine…

Incremental LearningSemantic Segmentation

Multi-branch Attentive Transformer

2020-06-18 · Yang Fan, Shufang Xie, Yingce Xia, Lijun Wu 외

While the multi-branch architecture is one of the key ingredients to the success of computer vision tasks, it has not been well investigated in natural language processing, especially sequence learning tasks. In this wor…

Code GenerationMachine TranslationNatural Language UnderstandingTranslation

Parameter Efficient Continual Learning for Sparse Event-Based Transformers

2026-08-27 · Vaishnavi Nagabhushana, Kartikay Agrawal, Ayon Borthakur arxiv

Robotic and edge intelligence systems operate in dynamic environments where data arrives continuously, requiring models to adapt while preserving previously learned knowledge under strict memory and energy constraints. W…

parameter-efficient fine-tuningclass-incremental learningContinual LearningEvent-based vision