paper-with-me

홈 › Papers

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models

2025-06-26 · Aristeidis Tsaris, Isaac Lyngaas, John Lagregren, Mohamed Wahib, Larry York, Prasanna Balaprakash, Dan Lu, Feiyi Wang, Xiao Wang

Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources such as varying physical groundings or data acquisition systems and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.

📄 PDF Abstract BibTeX arXiv:2506.21411

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiencyscientific discoveryWeather Forecasting

Methods 이 논문이 사용한 방법론

Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

Hierarchical Over-the-Air FedGradNorm

2022-12-14 · Cemil Vahapoglu, Matin Mortaheb, Sennur Ulukus

Multi-task learning (MTL) is a learning paradigm to learn multiple related tasks simultaneously with a single shared network where each task has a distinct personalized header network for fine-tuning. MTL can be integrat…

Federated LearningMulti-Task LearningPersonalized Federated Learning

Performance and Energy Trade-Off Analysis of Hierarchical Federated Learning for Plant Disease Classification

2026-04-28 · Athanasios Papanikolaou, Athanasios Tziouvaras, Pavlos Stoikos, Apostolos Xenakis 외 arxiv

Early detection of plant diseases is critical for improving crop productivity, while it also facilitates the foundations of precision agriculture. Recent advances in distributed deep learning have enabled plant disease c…

Federated Learning

Hierarchical Spatial and Channel Aggregation for Cross-domain Few-shot Segmentation

2026-06-23 · Sujun Sun, Mingwu Ren, Haofeng Zhang arxiv

Cross-domain Few-shot Segmentation (CD-FSS) aims to learn generalizable segmentation capability from abundant annotated samples in the source domain, enabling accurate segmentation of novel classes in the target domain w…

Cross-Domain Few-Shot

A Hierarchical Ensemble Pipeline for Anomaly Detection in ESA Satellite Telemetry

2026-04-22 · Lorenzo Riccardo Allegrini, Geremia Pompei arxiv

A hierarchical ensemble pipeline is introduced to address anomaly detection in multivariate telemetry data provided by European Space Agency (ESA). The method integrates shapelet-based and statistical feature extraction,…

Anomaly Detection

OSF: On Pre-training and Scaling of Sleep Foundation Models

2026-02-27 · Zitao Shuai, Zongzhe Xu, David Yang, Wei Wang 외 arxiv

Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from substantial heterogeneity across recording devices and cohorts. There have been growing efforts to build general-purpose foundation m…