paper-with-me

홈 › Papers

Advancing Direct Convolution using Convolution Slicing Optimization and ISA Extensions

2023-03-08 · Victor Ferrari, Rafael Sousa, Marcio Pereira, João P. L. de Carvalho, José Nelson Amaral, José Moreira, Guido Araujo

Convolution is one of the most computationally intensive operations that must be performed for machine-learning model inference. A traditional approach to compute convolutions is known as the Im2Col + BLAS method. This paper proposes SConv: a direct-convolution algorithm based on a MLIR/LLVM code-generation toolchain that can be integrated into machine-learning compilers . This algorithm introduces: (a) Convolution Slicing Analysis (CSA) - a convolution-specific 3D cache-blocking analysis pass that focuses on tile reuse over the cache hierarchy; (b) Convolution Slicing Optimization (CSO) - a code-generation pass that uses CSA to generate a tiled direct-convolution macro-kernel; and (c) Vector-Based Packing (VBP) - an architecture-specific optimized input-tensor packing solution based on vector-register shift instructions for convolutions with unitary stride. Experiments conducted on 393 convolutions from full ONNX-MLIR machine-learning models indicate that the elimination of the Im2Col transformation and the use of fast packing routines result in a total packing time reduction, on full model inference, of 2.0x - 3.9x on Intel x86 and 3.6x - 7.2x on IBM POWER10. The speed-up over an Im2Col + BLAS method based on current BLAS implementations for end-to-end machine-learning model inference is in the range of 9% - 25% for Intel x86 and 10% - 42% for IBM POWER10 architectures. The total convolution speedup for model inference is 12% - 27% on Intel x86 and 26% - 46% on IBM POWER10. SConv also outperforms BLAS GEMM, when computing pointwise convolutions, in more than 83% of the 219 tested instances.

📄 PDF Abstract BibTeX arXiv:2303.04739

Code (0)

등록된 구현이 없습니다.

Tasks

BlockingCode Generation

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Tensor Slicing and Optimization for Multicore NPUs

2023-04-06 · Rafael Sousa, Marcio Pereira, Yongin Kwon, TaeHo Kim 외

Although code generation for Convolution Neural Network (CNN) models has been extensively studied, performing efficient data slicing and parallelization for highly-constrai\-ned Multicore Neural Processor Units (NPUs) is…

Code GenerationCompiler Optimization

Revisiting Sliced Wasserstein on Images: From Vectorization to Convolution

2022-04-04 · Khai Nguyen, Nhat Ho

The conventional sliced Wasserstein is defined between two probability measures that have realizations as vectors. When comparing two probability measures over images, practitioners first need to vectorize images and the…

Designing by Training: Acceleration Neural Network for Fast High-Dimensional Convolution

2018-12-01 · NeurIPS 2018 12 · Longquan Dai, Liang Tang, Yuan Xie, Jinhui Tang

The high-dimensional convolution is widely used in various disciplines but has a serious performance problem due to its high computational complexity. Over the decades, people took a handmade approach to design fast algo…

Vocal Bursts Intensity Prediction

Real-time multi-view deconvolution

2015-03-27 · Benjamin Schmid, Jan Huisken

In light-sheet microscopy, overall image content and resolution are improved by acquiring and fusing multiple views of the sample from different directions. State-of-the-art multi-view (MV) deconvolution employs the poin…

GPU

Advancing RAN Slicing with Offline Reinforcement Learning

2023-12-16 · Kun Yang, Shu-ping Yeh, Menglei Zhang, Jerry Sydir 외

Dynamic radio resource management (RRM) in wireless networks presents significant challenges, particularly in the context of Radio Access Network (RAN) slicing. This technology, crucial for catering to varying user requi…

ManagementOffline RLreinforcement-learningReinforcement Learning+1