Greening AI Inference with Accuracy and Latency-aware User Incentives
The widespread use of AI services has raised concerns for its environmental sustainability, towards which recent studies have identified carbon emissions of AI inference as the major contributor. This paper introduces a framework for designing AI inference incentives based on the users' valuation for inference quality and latency, together with their environmental consciousness, while accounting for the tradeoff between carbon emissions and the two QoE parameters. Our approach can accommodate different tradeoffs, that depend on the size and complexity of the AI models and the allocation of resources to serve inference requests. The incentives can be offered through a practical two-tier service subscription that offers users a discount in exchange for reduced carbon emissions. The discounted service option gives the AI provider the flexibility to serve some percentage of inference requests at a lower quality and higher latency during periods of high carbon intensity.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Airport-Airline Coordination with Economic, Environmental and Social Considerations
In this paper, we examine the effect of various contracts between a socially concerned airport and an environmentally conscious airline regarding their profitability and channel coordination under two distinct settings. …
Dynamic Network Adaptation at Inference
Machine learning (ML) inference is a real-time workload that must comply with strict Service Level Objectives (SLOs), including latency and accuracy targets. Unfortunately, ensuring that SLOs are not violated in inferenc…
DiversityA sustainable development perspective on urban-scale roof greening priorities and benefits
Greenspaces are tightly linked to human well-being. Yet, rapid urbanization has exacerbated greenspace exposure inequality and declining human life quality. Roof greening has been recognized as an effective strategy to m…
Accuracy-Delay Trade-Off in LLM Offloading via Token-Level Uncertainty
Large language models (LLMs) offer significant potential for intelligent mobile services but are computationally intensive for resource-constrained devices. Mobile edge computing (MEC) allows such devices to offload infe…
LiveMind: Low-latency Large Language Models with Simultaneous Inference
In this paper, we introduce LiveMind, a novel low-latency inference framework for large language model (LLM) inference which enables LLMs to perform inferences with incomplete user input. By reallocating computational pr…
Collaborative InferenceLanguage ModelingLanguage ModellingLarge Language Model+1