paper-with-me

홈 › Papers

EdgeSight: Enabling Modeless and Cost-Efficient Inference at the Edge

2024-05-29 · ChonLam Lao, Jiaqi Gao, Ganesh Ananthanarayanan, Aditya Akella, Minlan Yu

Traditional ML inference is evolving toward modeless inference, which abstracts the complexity of model selection from users, allowing the system to automatically choose the most appropriate model for each request based on accuracy and resource requirements. While prior studies have focused on modeless inference within data centers, this paper tackles the pressing need for cost-efficient modeless inference at the edge -- particularly within its unique constraints of limited device memory, volatile network conditions, and restricted power consumption. To overcome these challenges, we propose EdgeSight, a system that provides cost-efficient EdgeSight serving for diverse DNNs at the edge. EdgeSight employs an edge-data center (edge-DC) architecture, utilizing confidence scaling to reduce the number of model options while meeting diverse accuracy requirements. Additionally, it supports lossy inference in volatile network environments. Our experimental results show that EdgeSight outperforms existing systems by up to 1.6x in P99 latency for modeless services. Furthermore, our FPGA prototype demonstrates similar performance at certain accuracy levels, with a power consumption reduction of up to 3.34x.

📄 PDF Abstract BibTeX arXiv:2405.19213

Code (0)

등록된 구현이 없습니다.

Tasks

Model Selection

Similar Papers 제목 키워드 기반

Spiker-LL: An Energy-Efficient FPGA Accelerator Enabling Adaptive Local Learning in Spiking Neural Networks

2026-05-18 · Alessio Caviglia, Filippo Marostica, Alessandro Savino, Stefano Di Carlo arxiv

Deploying adaptive intelligence at the edge remains challenging due to the high computational and energy cost of training neural models. Spiking Neural Networks (SNNs) offer a promising alternative, but enabling on-devic…

MEET: Mobility-Enhanced Edge inTelligence for Smart and Green 6G Networks

2022-10-27 · Yuxuan Sun, Bowen Xie, Sheng Zhou, Zhisheng Niu

Edge intelligence is an emerging paradigm for real-time training and inference at the wireless edge, thus enabling mission-critical applications. Accordingly, base stations (BSs) and edge servers (ESs) need to be densely…

Edge-InversionNet: Enabling Efficient Inference of InversionNet on Edge Devices

2023-10-14 · Zhepeng Wang, Isaacshubhanand Putla, Weiwen Jiang, Youzuo Lin

Seismic full waveform inversion (FWI) is a widely used technique in geophysics for inferring subsurface structures from seismic data. And InversionNet is one of the most successful data-driven machine learning models tha…

Geophysics

LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference

2026-07-30 · Feng Yang, Xinrui Ju, Keyang Zhang, Xiandong Meng 외 arxiv

Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud …

Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference

2025-07-11 · Chun-Ting Chen, HanGyeol Mun, Jian Meng, Mohamed S. Abdelfattah 외 arxiv

Edge inference for large language models (LLM) offers secure, low-latency, and cost-effective inference solutions. We emphasize that an edge accelerator should achieve high area efficiency and minimize external memory ac…