paper-with-me

홈 › Papers

Block-Parallel IDA* for GPUs (Extended Manuscript)

2017-05-08 · Satoru Horie, Alex Fukunaga

We investigate GPU-based parallelization of Iterative-Deepening A* (IDA*). We show that straightforward thread-based parallelization techniques which were previously proposed for massively parallel SIMD processors perform poorly due to warp divergence and load imbalance. We propose Block-Parallel IDA* (BPIDA*), which assigns the search of a subtree to a block (a group of threads with access to fast shared memory) rather than a thread. On the 15-puzzle, BPIDA* on a NVIDIA GRID K520 with 1536 CUDA cores achieves a speedup of 4.98 compared to a highly optimized sequential IDA* implementation on a Xeon E5-2670 core.

📄 PDF Abstract BibTeX arXiv:1705.02843

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

GPU accelerated matrix factorization of large scale data using block based approach

2023-01-02 · Prasad Bhavana, Vineet Padmanabhan

Matrix Factorization (MF) on large scale data takes substantial time on a Central Processing Unit (CPU). While Graphical Processing Unit (GPU)s could expedite the computation of MF, the available memory on a GPU is finit…

CPUGPU

Faster Multi-GPU Training with PPLL: A Pipeline Parallelism Framework Leveraging Local Learning

2024-11-19 · Xiuyuan Guo, Chengqi Xu, Guinan Guo, Feiyu Zhu 외

Currently, training large-scale deep learning models is typically achieved through parallel training across multiple GPUs. However, due to the inherent communication overhead and synchronization delays in traditional mod…

GPU

Minute-Long Videos with Dual Parallelisms

2025-05-27 · Zeqing Wang, Bowen Zheng, Xingyi Yang, Zhenxiong Tan 외

Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address this, we propose a novel distributed in…

DenoisingGPUVideo Generation

Pipe-BD: Pipelined Parallel Blockwise Distillation

2023-01-29 · Hongsun Jang, Jaewon Jung, Jaeyong Song, Joonsang Yu 외

Training large deep neural network models is highly challenging due to their tremendous computational and memory requirements. Blockwise distillation provides one promising method towards faster convergence by splitting …

GPU

Block Cascading: Training Free Acceleration of Block-Causal Video Models

2025-11-25 · Hmrishav Bandyopadhyay, Nikhil Pinnaparaju, Rahim Entezari, Jim Scott 외 arxiv

Block-causal video generation faces a stark speed-quality trade-off: small 1.3B models manage only 16 FPS while large 14B models crawl at 4.5 FPS, forcing users to choose between responsiveness and quality. Block Cascadi…

Video Generation