paper-with-me

홈 › Papers

Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach

2024-11-28 · Yijia Zhang, Zhihong Gou, Shijie Cao, Weigang Feng, Sicheng Zhang, Guohao Dai, Ningyi Xu

Deep Neural Networks (DNNs) have revolutionized various fields, but their deployment on GPUs often leads to significant energy consumption. Unlike existing methods for reducing GPU energy consumption, which are either hardware-inflexible or limited by workload constraints, this paper addresses the problem at the GPU kernel level. We propose a novel search-based compilation method to generate energy-efficient GPU kernels by incorporating energy efficiency into the search process. To accelerate the energy evaluation process, we develop an accurate energy cost model based on high-level kernel features. Furthermore, we introduce a dynamic updating strategy for the energy cost model, reducing the need for on-device energy measurements and accelerating the search process. Our evaluation demonstrates that the proposed approach can generate GPU kernels with up to 21.69% reduced energy consumption while maintaining low latency.

📄 PDF Abstract BibTeX arXiv:2411.18873

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

GreenAuto: An Automated Platform for Sustainable AI Model Design on Edge Devices

2025-01-25 · Xiaolong Tu, Dawei Chen, Kyungtae Han, Onur Altintas 외

We present GreenAuto, an end-to-end automated platform designed for sustainable AI model exploration, generation, deployment, and evaluation. GreenAuto employs a Pareto front-based search method within an expanded neural…

Neural Architecture Search

Towards Automated Kernel Generation in the Era of LLMs

2026-01-22 · Yang Yu, Peiyu Zang, Chi Hsu Tsai, Haiming Wu 외 arxiv

The performance of modern AI systems is fundamentally constrained by the quality of their underlying GPU kernels, which translate high-level algorithmic semantics into low-level hardware operations. Achieving near-optima…

LLM-Driven Kernel Evolution: Automating Driver Updates in Linux

2025-11-24 · Arina Kharlamova, Jiawen Liu, Tianyi Zhang, Xinrui Yang 외 arxiv

Linux kernel evolution breaks drivers through API/ABI changes, semantic shifts, and security-hardening updates. We introduce DRIVEBENCH, an executable corpus of kernel$\rightarrow$driver co-evolution cases, and AUTODRIVE…

Prompt Engineering

Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates for LLM-Based GPU Kernel Optimization

2026-06-24 · Jiading Gai, Shuai Zhang, Kaj Bostrom, Jin Huang 외 arxiv

We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating large language model (LLM) code generation with hardware profiler fe…

Code GenerationCode Search

FastKernels: Benchmarking GPU Kernel Generation in Production

2026-05-22 · Gabriele Oliaro, Yichao Fu, May Jiang, Owen Lu 외 arxiv

LLM-based agents for GPU kernel generation are advancing rapidly, yet their progress is fundamentally constrained by the benchmarks they optimize against. Existing benchmarks are poorly aligned with production inference …