paper-with-me

Papers

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

2026-07-08 · Jieying Wang, Shuyuan Fan, Mingkai Zheng, Zhao Zhang arxiv

Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the communication volume. Existing methods quantize gradients via linear or nonlinear mappings in Euclidean space, often degrading model performance because highly anisotropic gradients incur direction-dependent distortion. We present GIFT, a geometry-informed gradient scaling method that performs low-precision communication in geometry-aware coordinates. By transforming gradients into a near-isotropic space before quantization, GIFT makes low-precision representations substantially more faithful to their high-precision counterparts. GIFT only changes the coordinate system used for low-precision gradient communication and does not change the optimizer, training recipe, communication collective, or low-precision format. We also develop a simplified geometry-aware transformation algorithm with low-rank approximation and selective application to balance the computation overhead and communication reduction. We examine the empirical convergence of GIFT using Llama-300M and Llama-600M models. Our results show that GIFT reduces the end-to-end pretraining time of Llama-600M by 7.6% on 64 NVIDIA GH200 Superchips, while improving the downstream task preservation profile over direct Euclidean FP8 communication under the same optimizer and communication path.

📄 PDF Abstract BibTeX arXiv:2607.07494

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GIFT-SW: Gaussian noise Injected Fine-Tuning of Salient Weights for LLMs

2024-08-27 · Maxim Zhelnin, Viktor Moskvoretskii, Egor Shvetsov, Egor Venediktov 외

Parameter Efficient Fine-Tuning (PEFT) methods have gained popularity and democratized the usage of Large Language Models (LLMs). Recent studies have shown that a small subset of weights significantly impacts performance…

parameter-efficient fine-tuningQuantization

GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback

2026-03-28 · Giorgio Giannone, Anna Clare Doris, Amin Heyrani Nobari, Kai Xu 외 arxiv

Generating executable CAD programs from images requires alignment between visual geometry and symbolic program representations, a capability that current methods fail to learn reliably as design complexity increases. Exi…

Data AugmentationProgram Synthesis

Genomic Informational Field Theory (GIFT) to characterize genotypes involved in large phenotypic fluctuations

2023-07-05 · Cyril Rauch, Panagiota Kyratzi, Andras Paldi

Based on the normal distribution and its properties, i.e., average and variance, Fisher works have provided a conceptual framework to identify genotype-phenotype associations. While Fisher intuition has proved fruitful o…

GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention

2025-07-18 · Amro Abdalla, Ismail Shaheen, Dan DeGenaro, Rupayan Mallick 외 arxiv

We present GIFT: a {G}radient-aware {I}mmunization technique to defend diffusion models against malicious {F}ine-{T}uning while preserving their ability to generate safe content. Existing safety mechanisms like safety ch…

Gift Contagion in Online Groups: Evidence From Virtual Red Packets

2019-06-24 · Yuan Yuan, Tracy Liu, Chenhao Tan, Qian Chen 외

Gifts are important instruments for forming bonds in interpersonal relationships. Our study analyzes the phenomenon of gift contagion in online groups. Gift contagion encourages social bonds by prompting further gifts; i…

Experimental DesignMarketing