paper-with-me

Papers

CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference

2026-05-12 · Nan Xue, Shengkang Chen, Zhiyong Chen, Jiangchao Yao, Yaping Sun, Zixia Hu, Meixia Tao arxiv

As large language models (LLMs) move from centralized clouds to mobile edge environments, efficient serving must balance latency, energy consumption, and accuracy under constrained device-edge resources. Query-level routing between lightweight on-device models and stronger edge models provides a flexible mechanism to navigate this trade-off. However, existing routers are designed for centralized cloud settings and optimize token-level costs, failing to capture the dynamic latency and energy overheads in wireless edge deployments. In this paper, we formulate mobile edge LLM routing as a deployment-constrained, cost-aware decision problem, and propose CR^2, a two-stage device-edge routing framework. CR^2 decouples a lightweight on-device margin gate from an edge-side utility selector for deferred queries. The margin gate operates on frozen query embeddings and a user-specified cost weight to predict whether local execution is utility-optimal relative to the best edge alternative under the target operating point. We further introduce a conformal risk control (CRC) calibration procedure that maps each operating point to an acceptance threshold, enabling explicit control of the marginal false-acceptance risk under the full-information utility reference. Experiments on the routing task show that CR^2 closely matches a full-information reference router using only device-side signals before deferral. Compared with strong query-level baselines, CR^2 consistently improves the deployable accuracy-cost Pareto frontier and reduces normalized deployment cost by up to 16.9% at matched accuracy.

📄 PDF Abstract BibTeX arXiv:2605.12001

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dynamic Quality-Latency Aware Routing for LLM Inference in Wireless Edge-Device Networks

2025-08-15 · Rui Bao, Nan Xue, Yaping Sun, Zhiyong Chen arxiv

The integration of wireless communications and Large Language Models (LLMs) is poised to unlock ubiquitous intelligent services, yet deploying them in wireless edge-device collaborative environments presents a critical t…

Energy Optimized Congestion Control-Based Temperature Aware Routing Algorithm for Software Defined Wireless Body Area Networks

2020-02-27 · journal 2020 2 · Omar Ahmed, Fuji Ren, Ammar Hawbani, Yaser Al-Sharabi

Wireless Body Area Network (WBAN) is a promising cost-effective technology for the privacy confined military applications and healthcare applications like remote health monitoring, telemedicine, and e-health services. Th…

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

2026-08-28 · Ji'an Lei, Jian Huang arxiv

Agents with smaller language-model backbones are less expensive but can drift into persistent failure modes, whereas those with larger backbones are generally more reliable but more costly. This reliability-cost trade-of…

Proactive Routing to Interpretable Surrogates with Distribution-Free Safety Guarantees

2026-03-15 · Iqtedar Uddin, Mazin Khider, André Bauer arxiv

Model routing determines whether to use an accurate black-box model or a simpler surrogate that approximates it at lower cost or greater interpretability. In deployment settings, practitioners often wish to restrict surr…

Asynchronous Risk-Aware Multi-Agent Packet Routing for Ultra-Dense LEO Satellite Networks

2025-10-31 · Ke He, Thang X. Vu, Le He, Lisheng Fan 외 arxiv

The rise of ultra-dense LEO constellations creates a complex and asynchronous network environment, driven by their massive scale, dynamic topologies, and significant delays. This unique complexity demands an adaptive pac…