paper-with-me

홈 › Papers

AdapterTune: Zero-Initialized Low-Rank Adapters for Frozen Vision Transformers

2026-03-16 · Salim Khazem arxiv

Frozen-backbone transfer with Vision Transformers faces two under-addressed issues: optimization instability when adapters are naively inserted into a fixed feature extractor, and the absence of principled guidance for setting adapter capacity. We introduce AdapterTune, which augments each transformer block with a residual low-rank bottleneck whose up-projection is zero-initialized, guaranteeing that the adapted network starts exactly at the pretrained function and eliminates early-epoch representation drift. On the analytical side, we formalize adapter rank as a capacity budget for approximating downstream task shifts in feature space. The resulting excess-risk decomposition predicts monotonic but diminishing accuracy gains with increasing rank, an ``elbow'' behavior we confirm through controlled sweeps. We evaluate on 9 datasets and 3 backbone scales with multi-seed reporting throughout. On a core 5 dataset transfer suite, AdapterTune improves top-1 accuracy over head-only transfer by +14.9 points on average while training only 0.92 of the parameters required by full fine-tuning, and outperforms full fine-tuning on 10 of 15 dataset-backbone pairs. Across the full benchmark, AdapterTune improves over head-only transfer on every dataset-backbone pair tested. Ablations on rank, placement, and initialization isolate each design choice. The code is available at: https://github.com/salimkhazem/adaptertune

📄 PDF Abstract BibTeX arXiv:2603.14706

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReCoLoRA: Spectrum-Aware Recursive Consolidation for Continual LLM Fine-Tuning

2026-07-04 · Wentao Lu arxiv

Parameter-efficient fine-tuning adapts a large language model to one task cheaply, but across a task sequence LoRA-style methods keep stacking low-rank updates on the same frozen weight, so each new task tends to overwri…

parameter-efficient fine-tuning

SteeringDiffusion: A Bottlenecked Activation Control Interface for Diffusion Models

2026-05-03 · Fangzheng Wu, Brian Summa arxiv

We introduce SteeringDiffusion, a bottlenecked activation-level control interface for diffusion models that exposes a smooth, monotonic, and runtime-adjustable control surface over the content--style trade-off. Our metho…

TeRA: Vector-based Random Tensor Network for High-Rank Adaptation of Large Language Models

2025-09-03 · Yuxuan Gu, Wuyang Zhou, Giorgos Iacovides, Danilo Mandic arxiv

Parameter-Efficient Fine-Tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), have significantly reduced the number of trainable parameters needed in fine-tuning large language models (LLMs). The developments of Lo…

parameter-efficient fine-tuning

Random Initialization of Gated Sparse Adapters

2025-11-03 · Vi Retault, Yohaï-Eliel Berreby arxiv

When fine-tuning language models on new tasks, catastrophic forgetting -- performance degradation on previously-learned tasks -- is a ubiquitous problem. While Parameter-Efficient Fine-Tuning (PEFT) methods like LoRA add…

parameter-efficient fine-tuning

A Little Rank Goes a Long Way: Random Scaffolds with LoRA Adapters Are All You Need

2026-04-09 · Hananel Hazan, Yanbo Zhang, Benedikt Hartl, Michael Levin arxiv

How many of a neural network's parameters actually encode task-specific information? We investigate this question with LottaLoRA, a training paradigm in which every backbone weight is drawn at random and frozen; only low…