paper-with-me

홈 › Papers

Small Language Models as Compiler Experts: Auto-Parallelization for Heterogeneous Systems

2025-12-22 · Prathamesh Devadiga arxiv

Traditional auto-parallelizing compilers, reliant on rigid heuristics, struggle with the complexity of modern heterogeneous systems. This paper presents a comprehensive evaluation of small (approximately 1B parameter) language-model-driven compiler auto-parallelization. We evaluate three models: gemma3, llama3.2, and qwen2.5, using six reasoning strategies across 11 real-world kernels drawn from scientific computing, graph algorithms, and machine learning. Our system is benchmarked against strong compiler baselines, including LLVM Polly, TVM, and Triton. Across 376 total evaluations, the proposed approach achieves an average speedup of 6.81x and a peak performance of 43.25x on convolution operations. We analyze scalability, verify correctness using multiple sanitizers, and confirm robustness across diverse compilers and hardware platforms. Our results demonstrate that small, efficient language models can serve as powerful reasoning engines for complex compiler optimization tasks.

📄 PDF Abstract BibTeX arXiv:2512.19250

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning to Parallelize in a Shared-Memory Environment with Transformers

2022-04-27 · Re'em Harel, Yuval Pinter, Gal Oren

In past years, the world has switched to many-core and multi-core shared memory architectures. As a result, there is a growing need to utilize these architectures by introducing shared memory parallelization schemes to s…

Management

Advising OpenMP Parallelization via a Graph-Based Approach with Transformers

2023-05-16 · Tal Kadosh, Nadav Schneider, Niranjan Hasabnis, Timothy Mattson 외

There is an ever-present need for shared memory parallelization schemes to exploit the full potential of multi-core architectures. The most common parallelization API addressing this need today is OpenMP. Nevertheless, w…

Data Augmentation

OMPar: Automatic Parallelization with AI-Driven Source-to-Source Compilation

2024-09-23 · Tal Kadosh, Niranjan Hasabnis, Prema Soundararajan, Vy A. Vo 외

Manual parallelization of code remains a significant challenge due to the complexities of modern software systems and the widespread adoption of multi-core architectures. This paper introduces OMPar, an AI-driven tool de…

C++ code

Towards Neural Decompilation

2019-05-20 · Omer Katz, Yuval Olshaker, Yoav Goldberg, Eran Yahav

We address the problem of automatic decompilation, converting a program in low-level representation back to a higher-level human-readable programming language. The problem of decompilation is extremely important for secu…

C++ codeMachine TranslationTranslation

Generating GPU Compiler Heuristics using Reinforcement Learning

2021-11-23 · Ian Colbert, Jake Daly, Norm Rubin

GPU compilers are complex software programs with many optimizations specific to target hardware. These optimizations are often controlled by heuristics hand-designed by compiler experts using time- and resource-intensive…

Deep Reinforcement LearningGPUreinforcement-learningReinforcement Learning+1