paper-with-me

홈 › Papers

Joint Continual Learning of Local Language Models and Cloud Offloading Decisions with Budget Constraints

2026-01-29 · Evan Chen, Wenzhi Fang, Shiqiang Wang, Christopher Brinton arxiv

Locally deployed Small Language Models (SLMs) must continually support diverse tasks under strict memory and computation constraints, making selective reliance on cloud Large Language Models (LLMs) unavoidable. Regulating cloud assistance during continual learning is challenging, as naive reward-based reinforcement learning often yields unstable offloading behavior and exacerbates catastrophic forgetting as task distributions shift. We propose DA-GRPO, a dual-advantage extension of Group Relative Policy Optimization that incorporates cloud-usage constraints directly into advantage computation, avoiding fixed reward shaping and external routing models. This design enables the local model to jointly learn task competence and collaboration behavior, allowing cloud requests to emerge naturally during post-training while respecting a prescribed assistance budget. Experiments on mathematical reasoning and code generation benchmarks show that DA-GRPO improves post-switch accuracy, substantially reduces forgetting, and maintains stable cloud usage compared to prior collaborative and routing-based approaches.

📄 PDF Abstract BibTeX arXiv:2602.00166

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningMathematical ReasoningContinual LearningCode Generation

Similar Papers 제목 키워드 기반

Secure Computation Offloading in Blockchain based IoT Networks with Deep Reinforcement Learning

2019-08-15 · Dinh C. Nguyen, Pubudu N. Pathirana, Ming Ding, Aruna Seneviratne

For current and future Internet of Things (IoT) networks, mobile edge-cloud computation offloading (MECCO) has been regarded as a promising means to support delay-sensitive IoT applications. However, offloading mobile ta…

Deep Reinforcement LearningManagementreinforcement-learningReinforcement Learning (RL)

PECCO: A Profit and Cost-oriented Computation Offloading Scheme in Edge-Cloud Environment with Improved Moth-flame Optimisation

2022-08-09 · Jiashu Wu, Hao Dai, Yang Wang, Shigen Shen 외

With the fast growing quantity of data generated by smart devices and the exponential surge of processing demand in the Internet of Things (IoT) era, the resource-rich cloud centres have been utilised to tackle these cha…

Local-Cloud Inference Offloading for LLMs in Multi-Modal, Multi-Task, Multi-Dialogue Settings

2025-02-16 · Liangqi Yuan, Dong-Jun Han, Shiqiang Wang, Christopher G. Brinton

Compared to traditional machine learning models, recent large language models (LLMs) can exhibit multi-task-solving capabilities through multiple dialogues and multi-modal data sources. These unique characteristics of LL…

Network Offloading Policies for Cloud Robotics: a Learning-based Approach

2019-02-15 · Sandeep Chinchali, Apoorva Sharma, James Harrison, Amine Elhafsi 외

Today's robotic systems are increasingly turning to computationally expensive models such as deep neural networks (DNNs) for tasks like localization, perception, planning, and object detection. However, resource-constrai…

Decision MakingDeep Reinforcement Learningobject-detectionObject Detection+2

Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training

2025-09-28 · Wenzhi Fang, Dong-Jun Han, Liangqi Yuan, Evan Chen 외 arxiv

Device-cloud collaboration holds promise for deploying large language models (LLMs), leveraging lightweight on-device models for efficiency while relying on powerful cloud models for superior reasoning. A central challen…

Reinforcement Learning