paper-with-me

홈 › Papers

FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern

2025-08-30 · Ao Shen, Rui Zhang, Junping Zhao arxiv

As large language models (LLMs) continue to scale, multi-node deployment has become a necessity. Consequently, communication has become a critical performance bottleneck. Current intra-node communication libraries, like NCCL, typically make use of a single interconnect such as NVLink. This approach creates performance ceilings, especially on hardware like the H800 GPU where the primary interconnect's bandwidth can become a bottleneck, and leaves other hardware resources like PCIe and Remote Direct Memory Access (RDMA)-capable Network Interface Cards (NICs) largely idle during intensive workloads. We propose FlexLink, the first collective communication framework to the best of our knowledge designed to systematically address this by aggregating these heterogeneous links-NVLink, PCIe, and RDMA NICs-into a single, high-performance communication fabric. FlexLink employs an effective two-stage adaptive load balancing strategy that dynamically partitions communication traffic across all available links, ensuring that faster interconnects are not throttled by slower ones. On an 8-GPU H800 server, our design improves the bandwidth of collective operators such as AllReduce and AllGather by up to 26% and 27% over the NCCL baseline, respectively. This gain is achieved by offloading 2-22% of the total communication traffic to the previously underutilized PCIe and RDMA NICs. FlexLink provides these improvements as a lossless, drop-in replacement compatible with the NCCL API, ensuring easy adoption.

📄 PDF Abstract BibTeX arXiv:2510.15882

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System

2025-08-17 · Yunhua Fang, Rui Xie, Asad Ul Haq, Linsen Ma 외 arxiv

Large Language Model (LLM) inference is increasingly constrained by memory bandwidth, with frequent access to the key-value (KV) cache dominating data movement. While attention sparsity reduces some memory traffic, the r…

Data-parallel distributed training of very large models beyond GPU capacity

2018-11-29 · Samuel Matzek, Max Grossman, Minsik Cho, Anar Yusifov 외

GPUs have limited memory and it is difficult to train wide and/or deep models that cause the training process to go out of memory. It is shown in this paper how an open source tool called Large Model Support (LMS) can ut…

CPUGPU

Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity

2026-03-13 · Donglin Yu arxiv

Multimodal large language model (MLLM) inference splits into two phases with opposing hardware demands: vision encoding is compute-bound, while language generation is memory-bandwidth-bound. We show that under standard t…

Coherence boosting: When your pretrained language model is not paying enough attention

2021-10-15 · ACL 2022 5 · Nikolay Malkin, Zhen Wang, Nebojsa Jojic

Long-range semantic coherence remains a challenge in automatic language generation and understanding. We demonstrate that large language models have insufficiently learned the effect of distant words on next-token predic…

Language ModelingLanguage ModellingText Generation

From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples

2024-04-11 · Robert Vacareanu, Vlad-Andrei Negru, Vasile Suciu, Mihai Surdeanu

We analyze how well pre-trained large language models (e.g., Llama2, GPT-4, Claude 3, etc) can do linear and non-linear regression when given in-context examples, without any additional training or gradient updates. Our …

Language ModelingLanguage ModellingLarge Language Modelregression