paper-with-me

홈 › Papers

Adaptive Compression-Aware Split Learning and Inference for Enhanced Network Efficiency

2023-11-09 · Akrit Mudvari, Antero Vainio, Iason Ofeidis, Sasu Tarkoma, Leandros Tassiulas

The growing number of AI-driven applications in mobile devices has led to solutions that integrate deep learning models with the available edge-cloud resources. Due to multiple benefits such as reduction in on-device energy consumption, improved latency, improved network usage, and certain privacy improvements, split learning, where deep learning models are split away from the mobile device and computed in a distributed manner, has become an extensively explored topic. Incorporating compression-aware methods (where learning adapts to compression level of the communicated data) has made split learning even more advantageous. This method could even offer a viable alternative to traditional methods, such as federated learning techniques. In this work, we develop an adaptive compression-aware split learning method ('deprune') to improve and train deep learning models so that they are much more network-efficient, which would make them ideal to deploy in weaker devices with the help of edge-cloud resources. This method is also extended ('prune') to very quickly train deep learning models through a transfer learning approach, which trades off little accuracy for much more network-efficient inference abilities. We show that the 'deprune' method can reduce network usage by 4x when compared with a split-learning approach (that does not use our method) without loss of accuracy, while also improving accuracy over compression-aware split-learning by 4 percent. Lastly, we show that the 'prune' method can reduce the training time for certain models by up to 6x without affecting the accuracy when compared against a compression-aware split-learning approach.

📄 PDF Abstract BibTeX arXiv:2311.05739

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningFederated LearningTransfer Learning

Similar Papers 제목 키워드 기반

Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing

2025-11-06 · Mingyu Sung, Vikas Palakonda, Suhwan Im, Sunghwan Moon 외 arxiv

Large language models (LLMs) have achieved near-human performance across diverse reasoning tasks, yet their deployment on resource-constrained Internet-of-Things (IoT) devices remains impractical due to massive parameter…

DynSplit-KV: Dynamic Semantic Splitting for KVCache Compression in Efficient Long-Context LLM Inference

2026-02-03 · Jiancai Ye, Jun Liu, Qingchen Li, Tianlang Zhao 외 arxiv

Although Key-Value (KV) Cache is essential for efficient large language models (LLMs) inference, its growing memory footprint in long-context scenarios poses a significant bottleneck, making KVCache compression crucial. …

Channel Pruning In Quantization-aware Training: An Adaptive Projection-gradient Descent-shrinkage-splitting Method

2022-04-09 · Zhijian Li, Jack Xin

We propose an adaptive projection-gradient descent-shrinkage-splitting method (APGDSSM) to integrate penalty based channel pruning into quantization-aware training (QAT). APGDSSM concurrently searches weights in both the…

Quantization

NSC-SL: A Bandwidth-Aware Neural Subspace Compression for Communication-Efficient Split Learning

2026-02-02 · Zhen Fang, Miao Yang, Zehang Lin, Zheng Lin 외 arxiv

The expanding scale of neural networks poses a major challenge for distributed machine learning, particularly under limited communication resources. While split learning (SL) alleviates client computational burden by dis…

AVERY: Intent-Driven Adaptive VLM Split Computing via Embodied Self-Awareness for Efficient Disaster Response Systems

2025-11-22 · Rajat Bhattacharjya, Sing-Yao Wu, Hyunwoo Oh, Chaewon Nam 외 arxiv

Unmanned Aerial Vehicles (UAVs) in disaster response require complex, queryable intelligence that onboard CNNs cannot provide. While Vision-Language Models (VLMs) offer this semantic reasoning, their high resource demand…

Image Compression