paper-with-me

Papers

Teleportation With Null Space Gradient Projection for Optimization Acceleration

2025-02-17 · Zihao Wu, Juncheng Dong, Ahmed Aloui, Vahid Tarokh

Optimization techniques have become increasingly critical due to the ever-growing model complexity and data scale. In particular, teleportation has emerged as a promising approach, which accelerates convergence of gradient descent-based methods by navigating within the loss invariant level set to identify parameters with advantageous geometric properties. Existing teleportation algorithms have primarily demonstrated their effectiveness in optimizing Multi-Layer Perceptrons (MLPs), but their extension to more advanced architectures, such as Convolutional Neural Networks (CNNs) and Transformers, remains challenging. Moreover, they often impose significant computational demands, limiting their applicability to complex architectures. To this end, we introduce an algorithm that projects the gradient of the teleportation objective function onto the input null space, effectively preserving the teleportation within the loss invariant level set and reducing computational cost. Our approach is readily generalizable from MLPs to CNNs, transformers, and potentially other advanced architectures. We validate the effectiveness of our algorithm across various benchmark datasets and optimizers, demonstrating its broad applicability.

📄 PDF Abstract BibTeX arXiv:2502.11362

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Level Set Teleportation: An Optimization Perspective

2024-03-05 · Aaron Mishkin, Alberto Bietti, Robert M. Gower

We study level set teleportation, an optimization routine which tries to accelerate gradient descent (GD) by maximizing the gradient norm over a level set of the objective. While teleportation intuitively speeds-up GD vi…

LEMMA

NPAT Null-Space Projected Adversarial Training Towards Zero Deterioration

2024-09-18 · Hanyi Hu, Qiao Han, Kui Chen, Yao Yang

To mitigate the susceptibility of neural networks to adversarial attacks, adversarial training has emerged as a prevalent and effective defense strategy. Intrinsically, this countermeasure incurs a trade-off, as it sacri…

Data Augmentation

Symmetry Teleportation for Accelerated Optimization

2022-05-21 · Bo Zhao, Nima Dehmamy, Robin Walters, Rose Yu

Existing gradient-based optimization methods update parameters locally, in a direction that minimizes the loss function. We study a different approach, symmetry teleportation, that allows parameters to travel a large dis…

Second-order methods

Neural Teleportation

2020-12-02 · Marco Armenta, Thierry Judge, Nathan Painchaud, Youssef Skandarani 외

In this paper, we explore a process called neural teleportation, a mathematical consequence of applying quiver representation theory to neural networks. Neural teleportation "teleports" a network to a new position in the…

Position

GNSP: Gradient Null Space Projection for Preserving Cross-Modal Alignment in VLMs Continual Learning

2025-07-26 · Tiantian Peng, Yuyang Liu, Shuo Yang, Qiuhe Hong 외 arxiv

Contrastive Language-Image Pretraining has demonstrated remarkable zero-shot generalization by aligning visual and textual modalities in a shared embedding space. However, when continuously fine-tuned on diverse tasks, C…

Zero-shot GeneralizationKnowledge DistillationCross-Modal RetrievalContinual Learning