paper-with-me

홈 › Papers

Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems

2024-06-28 · Yufan Xia, Giuseppe Maria Junior Barca

BLAS Level 3 operations are essential for scientific computing, but finding the optimal number of threads for multi-threaded implementations on modern multi-core systems is challenging. We present an extension to the Architecture and Data-Structure Aware Linear Algebra (ADSALA) library that uses machine learning to optimize the runtime of all BLAS Level 3 operations. Our method predicts the best number of threads for each operation based on the matrix dimensions and the system architecture. We test our method on two HPC platforms with Intel and AMD processors, using MKL and BLIS as baseline BLAS implementations. We achieve speedups of 1.5 to 3.0 for all operations, compared to using the maximum number of threads. We also analyze the runtime patterns of different BLAS operations and explain the sources of speedup. Our work shows the effectiveness and generality of the ADSALA approach for optimizing BLAS routines on modern multi-core systems.

📄 PDF Abstract BibTeX arXiv:2406.19621

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Library 설명 없음
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication

2026-01-14 · Yufan Xia, Marco De La Pierre, Amanda S. Barnard, Giuseppe Maria Junior Barca arxiv

The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the…

Wearable Tracking of Eye and Body Movements During Breaching Training: Towards Real-Time Blast Injury Monitoring

2025-05-14 · Jeremy P. Kemmerer, James R. Williamson, Joseph Kim, Elizabeth Halford 외

Repeated exposure to blast overpressure in occupational settings has been associated with changes in cognitive and psychological health, as well as deficits in neurosensory subsystems. In this work, we describe a wearabl…

mlpack 3: a fast, flexible machine learning library

2018-06-18 · Journal of Open Source Software 2018 6 · Ryan R. Curtin, Marcus Edel, Mikhail Lozhnikov, Yannis Mentekidis 외

In the past several years, the field of machine learning has seen an explosion of interest and excitement, with hundreds or thousands of algorithms developed for different tasks every year. But a primary problem faced by…

BenchmarkingBIG-bench Machine LearningClusteringGPU+1

KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning

2026-02-15 · Kris Shengjun Dong, Sahil Modi, Dima Nikiforov, Sana Damani 외 arxiv

Optimizing CUDA code across multiple generations of GPU architectures is challenging, as achieving peak performance requires an extensive exploration of an increasingly complex, hardware-specific optimization space. Trad…

Reinforcement Learning

Deep-Learning-Driven Prefetching for Far Memory

2025-05-31 · Yutong Huang, Zhiyuan Guo, Yiying Zhang

Modern software systems face increasing runtime performance demands, particularly in emerging architectures like far memory, where local-memory misses incur significant latency. While machine learning (ML) has proven eff…

Deep Learning