paper-with-me

Papers

PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers

2024-11-28 · Gwangoo Yeo, Jiin Kim, Yujeong Choi, Minsoo Rhu

NVIDIA's Multi-Instance GPU (MIG) is a feature that enables system designers to reconfigure one large GPU into multiple smaller GPU slices. This work characterizes this emerging GPU and evaluates its effectiveness in designing high-performance AI inference servers. Our study reveals that the data preprocessing stage of AI inference causes significant performance bottlenecks to MIG. To this end, we present PREBA, which is a hardware/software co-design targeting MIG inference servers. Our first proposition is an FPGA-based data preprocessing accelerator that unlocks the full potential of MIG with domain-specific acceleration of data preprocessing. The MIG inference server unleashed from preprocessing overheads is then augmented with our dynamic batching system that enables high-performance inference. PREBA is implemented end-to-end in real systems, providing a 3.7x improvement in throughput, 3.4x reduction in tail latency, 3.5x improvement in energy-efficiency, and 3.0x improvement in cost-efficiency.

📄 PDF Abstract BibTeX arXiv:2411.19114

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

PREBA: Surgical Duration Prediction via PCA-Weighted Retrieval-Augmented LLMs and Bayesian Averaging Aggregation

2026-02-27 · Wanyin Wu, Kanxue Li, Baosheng Yu, Haoyun Zhao 외 arxiv

Accurate prediction of surgical duration is pivotal for hospital resource management. Although recent supervised learning approaches-from machine learning (ML) to fine-tuned large language models (LLMs)-have shown strong…

HASCO: Towards Agile HArdware and Software CO-design for Tensor Computation

2021-05-04 · Qingcheng Xiao, Size Zheng, Bingzhe Wu, Pengcheng Xu 외

Tensor computations overwhelm traditional general-purpose computing devices due to the large amounts of data and operations of the computations. They call for a holistic solution composed of both hardware acceleration an…

Bayesian OptimizationQ-Learning

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench

2026-05-13 · Qingyun Zou, Feng Yu, Hongshi Tan, Bingsheng He 외 arxiv

We ask whether agentic AI systems built for software engineering transfer to realistic hardware engineering. Existing hardware LLM benchmarks isolate sub-tasks but none jointly requires repository navigation, hierarchy-a…

Beyond Moore's Law: Harnessing the Redshift of Generative AI with Effective Hardware-Software Co-Design

2025-04-09 · Amir Yazdanbakhsh

For decades, Moore's Law has served as a steadfast pillar in computer architecture and system design, promoting a clear abstraction between hardware and software. This traditional Moore's computing paradigm has deepened …

Philosophy

Learned Hardware/Software Co-Design of Neural Accelerators

2020-10-05 · Zhan Shi, Chirag Sakhuja, Milad Hashemi, Kevin Swersky 외

The use of deep learning has grown at an exponential rate, giving rise to numerous specialized hardware and software systems for deep learning. Because the design space of deep learning software stacks and hardware accel…

Bayesian OptimizationDeep Learning