paper-with-me

홈 › Papers

Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

2026-06-30 · Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham, Toan Tran arxiv

State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, offering substantial memory savings while retaining parameter efficiency. However, most existing state-based methods typically apply only per-block control updates, which limits inter-block information exchange and restricts representational adaptation. Meanwhile, prior mechanisms that enable cross-block communication often introduce considerable computational overhead, reducing their practicality for efficient fine-tuning. We introduce Mixture-of-Control (MoC), a lightweight fine-tuning framework that adaptively integrates local and global control signals to enhance representation learning. MoC treats block-wise control states as experts in a sparse mixture-of-experts process, enabling efficient communication across transformer blocks. Empirical results across diverse transformer-based benchmarks demonstrate that MoC outperforms state-based methods while maintaining a comparable memory and computational efficiency.

📄 PDF Abstract BibTeX arXiv:2606.31397

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyRepresentation Learning

Similar Papers 제목 키워드 기반

MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment

2026-03-19 · Yipu Dou, Wang Yang arxiv

We study how to allocate a fixed supervised fine-tuning budget when three objectives must be balanced at once: multi-turn safety alignment, low over-refusal on benign boundary queries, and instruction following under ver…

Instruction Following

Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation

2026-07-03 · Wei Zhang, Lin Tang, Ming Zhao, Yuxuan Wang arxiv

Fine-tuning a single low-rank adapter on many domains at once is multi-task learning: the domains must be co-learned, and how they share the adapter decides whether they help or hurt one another. Most efficient fine-tuni…

Multi-Task Learning

MixPHM: Redundancy-Aware Parameter-Efficient Tuning for Low-Resource Visual Question Answering

2023-03-02 · CVPR 2023 1 · Jingjing Jiang, Nanning Zheng

Recently, finetuning pretrained Vision-Language Models (VLMs) has been a prevailing paradigm for achieving state-of-the-art performance in Visual Question Answering (VQA). However, as VLMs scale, finetuning full model pa…

Mixture-of-ExpertsQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Merge to Mix: Mixing Datasets via Model Merging

2025-05-21 · Zhixu Silvia Tao, Kasper Vinken, Hao-Wei Yeh, Avi Cooper 외

Mixing datasets for fine-tuning large models (LMs) has become critical for maximizing performance on downstream tasks. However, composing effective dataset mixtures typically relies on heuristics and trial-and-error, oft…

HFedMoE: Resource-aware Heterogeneous Federated Learning with Mixture-of-Experts

2026-01-02 · Zihan Fang, Zheng Lin, Senkang Hu, Yanan Ma 외 arxiv

While federated learning (FL) enables fine-tuning of large language models (LLMs) without compromising data privacy, the substantial size of an LLM renders on-device training impractical for resource-constrained clients,…

Federated Learning