paper-with-me

Papers

Resource Management for Low-latency Cooperative Fine-tuning of Foundation Models at the Network Edge

2024-07-13 · Hai Wu, Xu Chen, Kaibin Huang

The emergence of large-scale foundation models (FoMo's) that can perform human-like intelligence motivates their deployment at the network edge for devices to access state-of-the-art artificial intelligence. For better user experiences, the pre-trained FoMo's need to be adapted to specialized downstream tasks through fine-tuning techniques. To transcend a single device's memory and computation limitations, we advocate multi-device cooperation within the device-edge cooperative fine-tuning (DEFT) paradigm, where edge devices cooperate to simultaneously optimize different parts of fine-tuning parameters within a FoMo. However, the parameter blocks reside at different depths within a FoMo architecture, leading to varied computation latency-and-memory cost due to gradient backpropagation-based calculations. The heterogeneous on-device computation and memory capacities and channel conditions necessitate an integrated communication-and-computation allocation of local computation loads and communication resources to achieve low-latency (LoLa) DEFT. To this end, we consider the depth-ware DEFT block allocation problem. The involved optimal block-device matching is tackled by the proposed low-complexity Cutting-RecoUNting-CHecking (CRUNCH) algorithm, which is designed by exploiting the monotone-increasing property between block depth and computation latency-and-memory cost. Next, the joint bandwidth-and-block allocation makes the problem more sophisticated. We observe a splittable Lagrangian expression through the transformation and analysis of the original problem, where the variables indicating device involvement are introduced. Then, the dual ascent method is employed to tackle this problem iteratively. Through extensive experiments conducted on the GLUE benchmark, our results demonstrate significant latency reduction achievable by LoLa DEFT for fine-tuning a RoBERTa model.

📄 PDF Abstract BibTeX arXiv:2407.09873

Code (0)

등록된 구현이 없습니다.

Tasks

Management

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Residual Connection 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Adam 설명 없음

Similar Papers 제목 키워드 기반

Privacy-Aware Agent Collaboration for Dynamic VR Slice Management in 6G SD-RAN

2026-05-28 · Khaled M. Naguib, Soumaya Cherkaoui, Mahmoud M. Elmesalawy, Ahmed M. Abd El-Haleem 외 arxiv

Ultra-low latency and high throughput are required for Virtual Reality (VR) services in 6G networks, which presents critical challenges for Software-Defined Radio Access Networks (SD-RANs) dynamic resource management. Th…

Multi-agent Reinforcement Learning

Multi-Agentic AI for Fairness-Aware and Accelerated Multi-modal Large Model Inference in Real-world Mobile Edge Networks

2026-02-06 · Haiyuan Li, Hari Madhukumar, Shuangyi Yan, Yulei Wu 외 arxiv

Generative AI (GenAI) has transformed applications in natural language processing and content creation, yet centralized inference remains hindered by high latency, limited customizability, and privacy concerns. Deploying…

Non-Cooperative Resource Management for Intelligent Reflecting Surface Aided Networks

2022-09-01 · Wenhao Cai, Ming Li, Qian Liu

Intelligent reflecting surface (IRS) has emerged as a promising and revolutionizing technology for future wireless networks. Most existing IRS studies focus on simple cooperative systems which usually have a single frequ…

Management

The Subtle Art of Defection: Understanding Uncooperative Behaviors in LLM based Multi-Agent Systems

2025-11-19 · Devang Kulshreshtha, Wanyu Du, Raghav Jain, Srikanth Doss 외 arxiv

This paper introduces a novel framework for simulating and analyzing how uncooperative behaviors can destabilize or collapse LLM-based multi-agent systems. Our framework includes two key components: (1) a game theory-bas…

Cooperative Multi-Type Multi-Agent Deep Reinforcement Learning for Resource Management in Space-Air-Ground Integrated Networks

2023-08-08 · Hengxi Zhang, Huaze Tang, Wenbo Ding, Xiao-Ping Zhang

The Space-Air-Ground Integrated Network (SAGIN), integrating heterogeneous devices including low earth orbit (LEO) satellites, unmanned aerial vehicles (UAVs), and ground users (GUs), holds significant promise for advanc…

Deep Reinforcement LearningManagement