paper-with-me

홈 › Papers

ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving

2026-05-30 · Seokjin Go, Marko Scrbak, Ephrem Wu, Srilatha Manne, Divya Mahajan arxiv

In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execution, where the slowest GPU determines layer latency. This performance variability is inherent to modern accelerators: manufacturing variation, power limits, and thermal conditions introduce measurable execution-time differences across nominally identical GPUs. The core challenge is that MoE execution-time imbalance arises from the interaction of workload skew and hardware asymmetry. Token routing produces uneven and layer-varying expert loads, while GPU throughput depends on device-specific operating characteristics and workload intensity. Prior work mitigates routing skew but assumes homogeneous hardware, optimizing token balance rather than execution latency. As a result, even balanced token assignments can leave hardware-induced stragglers unaddressed. Thus, we propose Variability-Informed Binning of Experts (ViBE), a hardware-aware expert placement framework that minimizes execution-time imbalance across GPUs. ViBE combines per-GPU performance modeling with expert activation profiling to assign high-load experts to faster devices and low-load experts to slower ones, reducing layer-level stragglers without modifying model semantics or hardware. Because both workload characteristics and effective GPU throughput can shift across serving conditions, ViBE supports lightweight recalibration under workload/performance drift to refresh its routing and performance estimates when needed. Results show that ViBE consistently reduces execution-time imbalance and improves SLO attainment by 14%, while lowering P90 TTFT by up to 45%. We further show that the impact of hardware variability increases at scale, making variability-aware placement important for efficient, high-utilization LLM serving.

📄 PDF Abstract BibTeX arXiv:2606.00735

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?

2026-05-07 · Keisuke Kamahori, Shihang Li, Simon Peter, Baris Kasikci arxiv

For years, we have built LLM serving systems like any other critical infrastructure: a single general-purpose stack, hand-tuned over many engineer-years, meant to support every model and workload. In this paper, we take …

Towards Efficient IMC Accelerator Design Through Joint Hardware-Workload Co-optimization

2024-10-22 · Olga Krestinskaya, Mohammed E. Fouda, Ahmed Eltawil, Khaled N. Salama

Designing generalized in-memory computing (IMC) hardware that efficiently supports a variety of workloads requires extensive design space exploration, which is infeasible to perform manually. Optimizing hardware individu…

Joint Hardware-Workload Co-Optimization for In-Memory Computing Accelerators

2026-03-04 · Olga Krestinskaya, Mohammed E. Fouda, Ahmed Eltawil, Khaled N. Salama arxiv

Software-hardware co-design is essential for optimizing in-memory computing (IMC) hardware accelerators for neural networks. However, most existing optimization frameworks target a single workload, leading to highly spec…

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

2025-11-01 · Ritik Raj, Souvik Kundu, Ishita Vohra, Hong Wang 외 arxiv

Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse task execution need, such serving heavily rely o…

Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration

2026-06-17 · Xhevahire Tërnava arxiv

In vibe coding, an emerging AI-driven paradigm, an LLM generates an entire program from a natural language prompt, but what happens to the variability that traditional software engineering carefully builds into code? To …