Cascade: A Platform for Delay-Sensitive Edge Intelligence
Interactive intelligent computing applications are increasingly prevalent, creating a need for AI/ML platforms optimized to reduce per-event latency while maintaining high throughput and efficient resource management. Yet many intelligent applications run on AI/ML platforms that optimize for high throughput even at the cost of high tail-latency. Cascade is a new AI/ML hosting platform intended to untangle this puzzle. Innovations include a legacy-friendly storage layer that moves data with minimal copying and a "fast path" that collocates data and computation to maximize responsiveness. Our evaluation shows that Cascade reduces latency by orders of magnitude with no loss of throughput.
Code (1)
Tasks
ManagementSimilar Papers 제목 키워드 기반
Smart Surveillance as an Edge Network Service: from Harr-Cascade, SVM to a Lightweight CNN
Edge computing efficiently extends the realm of information technology beyond the boundary defined by cloud computing paradigm. Performing computation near the source and destination, edge computing is promising to addre…
Cloud ComputingEdge-computingHuman Detectionobject-detection+1Caching and Computation Offloading in High Altitude Platform Station (HAPS) Assisted Intelligent Transportation Systems
Edge intelligence, a new paradigm to accelerate artificial intelligence (AI) applications by leveraging computing resources on the network edge, can be used to improve intelligent transportation systems (ITS). However, d…
Edge-computingEntire Space Cascade Delayed Feedback Modeling for Effective Conversion Rate Prediction
Conversion rate (CVR) prediction is an essential task for large-scale e-commerce platforms. However, refund behaviors frequently occur after conversion in online shopping systems, which drives us to pay attention to effe…
PredictionRecommendation SystemsSelection biasGPU Cluster Scheduling for Network-Sensitive Deep Learning
We propose a novel GPU-cluster scheduler for distributed DL (DDL) workloads that enables proximity based consolidation of GPU resources based on the DDL jobs' sensitivities to the anticipated communication-network delays…
Deep LearningGPUSchedulingModeling Cascaded Delay Feedback for Online Net Conversion Rate Prediction: Benchmark, Insights and Solutions
In industrial recommender systems, conversion rate (CVR) is widely used for traffic allocation, but it fails to fully reflect recommendation effectiveness because it ignores refund behavior. To better capture true user s…