Contextual Attention Modulation: Towards Efficient Multi-Task Adaptation in Large Language Models
Large Language Models (LLMs) possess remarkable generalization capabilities but struggle with multi-task adaptation, particularly in balancing knowledge retention with task-specific specialization. Conventional fine-tuning methods suffer from catastrophic forgetting and substantial resource consumption, while existing parameter-efficient methods perform suboptimally in complex multi-task scenarios. To address this, we propose Contextual Attention Modulation (CAM), a novel mechanism that dynamically modulates the representations of self-attention modules in LLMs. CAM enhances task-specific features while preserving general knowledge, thereby facilitating more effective and efficient adaptation. For effective multi-task adaptation, CAM is integrated into our Hybrid Contextual Attention Modulation (HyCAM) framework, which combines a shared, full-parameter CAM module with multiple specialized, lightweight CAM modules, enhanced by a dynamic routing strategy for adaptive knowledge fusion. Extensive experiments on heterogeneous tasks, including question answering, code generation, and logical reasoning, demonstrate that our approach significantly outperforms existing approaches, achieving an average performance improvement of 3.65%. The implemented code and data are available to ease reproducibility at https://github.com/Applied-Machine-Learning-Lab/HyCAM.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringGeneral KnowledgeLogical ReasoningCode GenerationSimilar Papers 제목 키워드 기반
Exploring Contextual Flux in Large Language Models: A Novel Approach to Self-Modulating Semantic Networks
Self-modulating mechanisms introduce dynamic adaptation capabilities within language models through contextual realignment strategies that influence token embedding trajectories across extended sequences. Contextual Flux…
Text GenerationNeuroLoRA: Context-Aware Neuromodulation for Parameter-Efficient Multi-Task Adaptation
Parameter-Efficient Fine-Tuning (PEFT) techniques, particularly Low-Rank Adaptation (LoRA), have become essential for adapting Large Language Models (LLMs) to downstream tasks. While the recent FlyLoRA framework successf…
parameter-efficient fine-tuningComputational EfficiencyContinual LearningSelf-supervised Learning with Speech Modulation Dropout
We show that training a multi-headed self-attention-based deep network to predict deleted, information-dense 2-8 Hz speech modulations over a 1.5-second section of a speech utterance is an effective way to make machines …
Automatic Speech RecognitionSelf-Supervised Learningspeech-recognitionSpeech RecognitionCellular neuromodulation in artificial networks
Animals excel at adapting their intentions, attention, and actions to the environment, making them remarkably efficient at interacting with a rich, unpredictable and ever-changing external world, a property that intellig…
Meta-LearningIntroducing Neuromodulation in Deep Neural Networks to Learn Adaptive Behaviours
Animals excel at adapting their intentions, attention, and actions to the environment, making them remarkably efficient at interacting with a rich, unpredictable and ever-changing external world, a property that intellig…
Meta Reinforcement LearningReinforcement Learning