paper-with-me

홈 › Papers

FedFrozen: Two-Stage Federated Optimization via Attention Kernel Freezing

2026-05-07 · Junye Du, Zhenghao Li, Yushi Feng, Long Feng arxiv

Federated learning with heterogeneous clients remains a significant challenge for deep learning, primarily due to client drift arising from inconsistent local updates. Existing federated optimization methods typically address this issue through objective-level regularization or update-correction mechanisms. Recent studies, however, suggest that Transformer-based architectures may be inherently more robust than conventional models under heterogeneous federated training. Motivated by this observation, we investigate how different parameter components within the attention mechanism influence federated optimization. Specifically, we decompose the attention module into a query/key block, which determines the attention kernel, and a value block, which performs semantic transformation under the induced kernel. Based on this perspective, we propose FedFrozen, a two-stage federated optimization framework that first performs full-model warm-up training and then freezes the query/key block while continuing to optimize the value block. Under a linear-attention formulation, we show that the warm-up stage can be interpreted as an inexact descent procedure on a regularized kernel-profile objective, while the frozen stage reduces to a restricted value-block optimization problem under a fixed attention kernel. Our analysis further reveals an explicit trade-off that governs the choice of warm-up length. Simulations validate the predicted bias-drift behavior, and real-data experiments demonstrate that FedFrozen improves both the stability and effectiveness of Transformer models in heterogeneous federated learning.

📄 PDF Abstract BibTeX arXiv:2605.06446

Code (0)

등록된 구현이 없습니다.

Tasks

Federated Learning

Similar Papers 제목 키워드 기반

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

2026-04-16 · Marcin Spoczynski, Daniel Fleischer, Moshe Berchansky, Gabriela Ben-Melech Stan 외 arxiv

Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization, memory access coalescing, tile size tuning, and architecture-specifi…

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

2026-08-17 · Joyjit Kundu, Ben Stoffelen, Kaili Wang, Peter Vrancx 외 arxiv

We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a de…

Federated Spectral Clustering via Secure Similarity Reconstruction

2023-09-21 · NeurIPS 2023 11

Federated learning has a significant advantage in protecting information privacy. Many scholars proposed various secure learning methods within the framework of federated learning but the study on secure federated unsupe…

TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels

2022-07-13 · Yaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma 외

State-of-the-art federated learning methods can perform far worse than their centralized counterparts when clients have dissimilar data distributions. For neural networks, even when centralized SGD easily finds a solutio…

Federated Learning

FedST: Secure Federated Shapelet Transformation for Time Series Classification

2023-02-21 · Zhiyu Liang, Hongzhi Wang

This paper explores how to build a shapelet-based time series classification (TSC) model in the federated learning (FL) scenario, that is, using more data from multiple owners without actually sharing the data. We propos…

ClassificationFederated LearningPrivacy PreservingTime Series+2