paper-with-me

홈 › Papers

Slide FFT on a homogeneous mesh in wafer-scale computing

2024-01-04 · Maurice H. P. M. van Putten, Leighton Wilson, Adam W. Lavely, Mark Hair

Searches for signals at low signal-to-noise ratios frequently involve the Fast Fourier Transform (FFT). For high-throughput searches, we here consider FFT on the homogeneous mesh of Processing Elements (PEs) of a wafer-scale engine (WSE). To minimize memory overhead in the inherently non-local FFT algorithm, we introduce a new synchronous slide operation ({\em Slide}) exploiting the fast interconnect between adjacent PEs. Feasibility of compute-limited performance is demonstrated in linear scaling of Slide execution times with varying array size in preliminary benchmarks on the CS-2 WSE. The proposed implementation appears opportune to accelerate and open the full discovery potential of FFT-based signal processing in multi-messenger astronomy.

📄 PDF Abstract BibTeX arXiv:2401.05427

Code (0)

등록된 구현이 없습니다.

Tasks

Astronomy

Similar Papers 제목 키워드 기반

WaferLLM: Large Language Model Inference at Wafer Scale

2025-02-06 · Congjie He, Yeqi Huang, Pei Mu, Ziming Miao 외

Emerging AI accelerators increasingly adopt wafer-scale manufacturing technologies, integrating hundreds of thousands of AI cores in a mesh architecture with large distributed on-chip memory (tens of GB in total) and ult…

GPULanguage ModelingLanguage ModellingLarge Language Model+1

DarwinWafer: A Wafer-Scale Neuromorphic Chip

2025-08-30 · Xiaolei Zhu, Xiaofei Jin, Ziyang Kang, Chonghui Sun 외 arxiv

Neuromorphic computing promises brain-like efficiency, yet today's multi-chip systems scale over PCBs and incur orders-of-magnitude penalties in bandwidth, latency, and energy, undermining biological algorithms and syste…

FRED: Flexible REduction-Distribution Interconnect and Communication Implementation for Wafer-Scale Distributed Training of DNN Models

2024-06-28 · Saeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 외

Distributed Deep Neural Network (DNN) training is a technique to reduce the training overhead by distributing the training tasks into multiple accelerators, according to a parallelization strategy. However, high-performa…

Full Wafer Redistribution and Wafer Embedding as Key Technologies for a Multi-Scale Neuromorphic Hardware Cluster

2018-01-15 · Kai Zoschke, Maurice Güttler, Lars Böttcher, Andreas Grübl 외

Together with the Kirchhoff-Institute for Physics(KIP) the Fraunhofer IZM has developed a full wafer redistribution and embedding technology as base for a large-scale neuromorphic hardware system. The paper will give an …

BrainScaleS Large Scale Spike Communication using Extoll

2021-11-30 · Tobias Thommes, Niels Buwen, Andreas Grübl, Eric Müller 외

The BrainScaleS Neuromorphic Computing System is currently connected to a compute cluster via Gigabit-Ethernet network technology. This is convenient for the currently used experiment mode, where neuronal networks cover …