paper-with-me

Papers

RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts

2026-04-28 · Vyom Sharma, Debajyoti Datta arxiv

The optimal kernel configuration for Mixture-of-Experts (MoE) inference depends on both batch size and the expert routing distribution, yet production systems dispatch from batch size alone, leaving 10-70% of kernel throughput unrealized. We present RaMP, a routing-aware dispatch framework. A performance-region analysis derives, from hardware constants alone, when each optimization helps, correctly predicting all 8 tested architectures, including 3 unseen. A four-parameter wave cost model selects the fastest configuration from the runtime expert histogram, achieving 0.93% mean regret versus exhaustive search, fitted from just 10-24 minutes of one-time profiling per model. Because the model depends only on CTA grid geometry, it is kernel-agnostic: applied to Alpha-MoE, it delivers 1.14x with no source modification. Paired with a co-designed CuTe DSL kernel exposing 134-268 polymorphic configurations, RaMP delivers 1.22x kernel speedup over static dispatch and 1.30x end-to-end speedup in vLLM serving over Triton, 1.41x over DeepGEMM, and 1.13x over FlashInfer CUTLASS.

📄 PDF Abstract BibTeX arXiv:2604.26039

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search for LLM Inference

2026-05-12 · Wenxin Dong, Mingqing Hu, Guanghui Yu, Qiang Fu 외 arxiv

When large language models (LLMs) serve real-time inference in commercial online advertising systems, end-to-end latency must be strictly bounded to the millisecond range. Yet every token generated during the decode phas…

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis

2026-06-08 · Jaber Jaber, Osama Jaber arxiv

AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA. The contribution is the…

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems

2026-05-26 · Yipeng Ouyang, Xin Huang, Bingjie Liu, Zhongchun Zheng 외 arxiv

LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain largely centered on static, isolated, and short-horizon benchmarks th…

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel

2026-04-14 · Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai 외 arxiv

Modern GPU workloads, especially large language model (LLM) inference, suffer from kernel launch overheads and coarse synchronization that limit inter-kernel parallelism. Recent megakernel techniques fuse multiple operat…

OWLOOP: Interfaces for Mapping OWL Axioms into OOP Hierarchies

2024-04-14 · Luca Buoncompagni, Fulvio Mastrogiovanni

The paper tackles the issue of mapping logic axioms formalised in the Ontology Web Language (OWL) within the Object-Oriented Programming (OOP) paradigm. The issues of mapping OWL axioms hierarchies and OOP objects hierar…