paper-with-me

Papers

RAPID: Redundancy-Aware and Compatibility-Optimal Edge-Cloud Partitioned Inference for Diverse VLA Models

2026-03-09 · Zihao Zheng, Sicheng Tian, Hangyu Cao, Chenyue Li, Jiayu Chen, Maoliang Li, Xinhao Sun, Hailong Zou, Guojie Luo, Xiang Chen arxiv

Vision Language Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) inference offers an effective fix by easing edge-device computing pressure to meet real-time needs. However, existing ECC frameworks are suboptimal for VLA models due to two challenges: (1) Mainstream environment-oriented edge-cloud partitioning methods are susceptible to interference from visual noise; (2) Existing edge-cloud partitioning methods overlook the step-wise redundancy unique to embodied tasks, thereby disrupting the physical continuity of motion. To address these issues, we propose a novel ECC inference framework, termed RAPID. Specifically, we developed an implementation tailored to the proposed framework. Experiments demonstrate this achieves a speedup of up to 1.73x with only 5%~7% overhead.

📄 PDF Abstract BibTeX arXiv:2603.07949

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training

2025-12-25 · Lei Liu, Hao Zhu, Yue Shen, Zhixuan Chu 외 arxiv

Continual Pre-training (CPT) serves as a fundamental approach for adapting foundation models to domain-specific applications. Scaling laws for pre-training define a power-law relationship between dataset size and the tes…

RAPID: Layer-Wise Redundancy-Aware Pruning and Importance-Driven Token Merging for Efficient ViT

2026-06-06 · Kyumin Choi, Ikbeom Jang arxiv

Vision Transformers (ViTs) achieve strong performance but suffer from high computational costs due to quadratic self-attention complexity. Although token reduction techniques such as pruning and merging mitigate this, th…

HoliTom: Holistic Token Merging for Fast Video Large Language Models

2025-05-27 · Kele Shao, Keda Tao, Can Qin, Haoxuan You 외

Video large language models (video LLMs) excel at video comprehension but face significant computational inefficiency due to redundant video tokens. Existing token pruning methods offer solutions. However, approaches ope…

ReLoRA: Knowledge-Reusing Adaptation for Fast Rollout of Evolving LLM Services

2026-05-23 · Yang Xu, Zihuai Xu, Hongli Xu, Yunming Liao 외 arxiv

Large Language Models (LLMs) are increasingly deployed as continuously evolving services, where frequent base-model updates may invalidate previously deployed task-specific Low-Rank Adaptation (LoRA) adapters. For servic…

Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation

2024-12-17 · Dongyue Wu, Zilin Guo, Li Yu, Nong Sang 외

In recent years, semantic segmentation has flourished in various applications. However, the high computational cost remains a significant challenge that hinders its further adoption. The filter pruning method for structu…

image-classificationImage ClassificationSegmentationSemantic Segmentation