paper-with-me

Papers

Don't Forget the Nonlinearity: Unlocking Activation Functions in Efficient Fine-Tuning

2025-09-16 · Bo Yin, Xingyi Yang, Xinchao Wang arxiv

Existing parameter-efficient fine-tuning (PEFT) methods primarily adapt weight matrices while keeping activation functions fixed. We introduce \textbf{NoRA}, the first PEFT framework that directly adapts nonlinear activation functions in pretrained transformer-based models. NoRA replaces fixed activations with learnable rational functions and applies structured low-rank updates to numerator and denominator coefficients, with a group-wise design that localizes adaptation and improves stability at minimal cost. On vision transformers trained on CIFAR-10 and CIFAR-100, NoRA matches or exceeds full fine-tuning while updating only 0.4\% of parameters (0.02M), achieving accuracy gains of +0.17\% and +0.27\%. When combined with LoRA (\textbf{NoRA++}), it outperforms LoRA and DoRA under matched training budgets by adding fewer trainable parameters. On LLaMA3-8B instruction tuning, NoRA++ consistently improves generation quality, yielding average MMLU gains of +0.3\%--0.8\%, including +1.6\% on STEM (Alpaca) and +1.3\% on OpenOrca. We further show that NoRA constrains adaptation to a low-dimensional functional subspace, implicitly regularizing update magnitude and direction. These results establish activation-space tuning as a complementary and highly parameter-efficient alternative to weight-based PEFT, positioning activation functions as first-class objects for model adaptation.

📄 PDF Abstract BibTeX arXiv:2509.13240

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

Nonlinearity Enhanced Adaptive Activation Functions

2024-03-29 · David Yevick

A general procedure for introducing parametric, learned, nonlinearity into activation functions is found to enhance the accuracy of representative neural networks without requiring significant additional computational re…

Unlocking the Power of Function Vectors for Characterizing and Mitigating Catastrophic Forgetting in Continual Instruction Tuning

2025-02-16 · Gangwei Jiang, Caigao Jiang, Zhaoyi Li, Siqiao Xue 외

Catastrophic forgetting (CF) poses a significant challenge in machine learning, where a model forgets previously learned information upon learning new tasks. Despite the advanced capabilities of Large Language Models (LL…

Continual Learning

Activation function dependence of the storage capacity of treelike neural networks

2020-07-21 · Jacob A. Zavatone-Veth, Cengiz Pehlevan

The expressive power of artificial neural networks crucially depends on the nonlinearity of their activation functions. Though a wide variety of nonlinear activation functions have been proposed for use in artificial neu…

Effect of shapes of activation functions on predictability in the echo state network

2019-05-22 · Hanten Chang, Shinji Nakaoka, Hiroyasu Ando

We investigate prediction accuracy for time series of Echo state networks with respect to several kinds of activation functions. As a result, we found that some kinds of activation functions with an appropriate nonlinear…

Time SeriesTime Series Analysis

Efficient Reinforcement Learning by Reducing Forgetting with Elephant Activation Functions

2025-09-23 · Qingfeng Lan, Gautham Vasan, A. Rupam Mahmood arxiv

Catastrophic forgetting has remained a significant challenge for efficient reinforcement learning for decades (Ring 1994, Rivest and Precup 2003). While recent works have proposed effective methods to mitigate this issue…

Reinforcement Learning