paper-with-me

홈 › Papers

Federating to Grow Transformers with Constrained Resources without Model Sharing

2024-06-19 · Shikun Shen, Yifei Zou, Yuan Yuan, Yanwei Zheng, Peng Li, Xiuzhen Cheng, Dongxiao Yu

The high resource consumption of large-scale models discourages resource-constrained users from developing their customized transformers. To this end, this paper considers a federated framework named Fed-Grow for multiple participants to cooperatively scale a transformer from their pre-trained small models. Under the Fed-Grow, a Dual-LiGO (Dual Linear Growth Operator) architecture is designed to help participants expand their pre-trained small models to a transformer. In Dual-LiGO, the Local-LiGO part is used to address the heterogeneity problem caused by the various pre-trained models, and the Global-LiGO part is shared to exchange the implicit knowledge from the pre-trained models, local data, and training process of participants. Instead of model sharing, only sharing the Global-LiGO strengthens the privacy of our approach. Compared with several state-of-the-art methods in simulation, our approach has higher accuracy, better precision, and lower resource consumption on computations and communications. To the best of our knowledge, most of the previous model-scaling works are centralized, and our work is the first one that cooperatively grows a transformer from multiple pre-trained heterogeneous models with the user privacy protected in terms of local data and models. We hope that our approach can extend the transformers to the broadly distributed scenarios and encourage more resource-constrained users to enjoy the bonus taken by the large-scale transformers.

📄 PDF Abstract BibTeX arXiv:2406.13450

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Federating Dynamic Models using Early-Exit Architectures for Automatic Speech Recognition on Heterogeneous Clients

2024-05-27 · Mohamed Nabih Ali, Alessio Brutti, Daniele Falavigna

Automatic speech recognition models require large amounts of speech recordings for training. However, the collection of such data often is cumbersome and leads to privacy concerns. Federated learning has been widely used…

Automatic Speech RecognitionFederated Learningspeech-recognitionSpeech Recognition

Chain-of-Thought Enhanced Shallow Transformers for Wireless Symbol Detection

2025-06-26 · Li Fan, Peng Wang, Jing Yang, Cong Shen

Transformers have shown potential in solving wireless communication problems, particularly via in-context learning (ICL), where models adapt to new tasks through prompts without requiring model updates. However, prior IC…

Computational EfficiencyIn-Context Learning

A Spatially Separable Attention Mechanism for massive MIMO CSI Feedback

2022-08-05 · Sharan Mourya, SaiDhiraj Amuru, Kiran Kumar Kuchi

Channel State Information (CSI) Feedback plays a crucial role in achieving higher gains through beamforming. However, for a massive MIMO system, this feedback overhead is huge and grows linearly with the number of antenn…

Compressive Sensing

How Merge-Tolerant Are Vision Transformers for Wheat Phenotyping?

2026-08-24 · Simon Ravé, Pejman Rasti, David Rousseau arxiv

Vision-based wheat phenotyping requires repeated measurements under deployment constraints, from growth-stage recognition to wheat-head counting and organ segmentation. Plain Vision Transformers (ViTs) provide a common a…

Head Detection

Mix-modal Federated Learning for MRI Image Segmentation

2025-09-02 · Guyue Hu, Siyuan Song, Jingpeng Sun, Zhe Jin 외 arxiv

Magnetic resonance imaging (MRI) image segmentation is crucial in diagnosing and treating many diseases, such as brain tumors. Existing MRI image segmentation methods mainly fall into a centralized multimodal paradigm, w…

Federated LearningImage Segmentation