paper-with-me

홈 › Papers

Group and Shuffle: Efficient Structured Orthogonal Parametrization

2024-06-14 · Mikhail Gorbunov, Nikolay Yudin, Vera Soboleva, Aibek Alanov, Alexey Naumov, Maxim Rakhuba

The increasing size of neural networks has led to a growing demand for methods of efficient fine-tuning. Recently, an orthogonal fine-tuning paradigm was introduced that uses orthogonal matrices for adapting the weights of a pretrained model. In this paper, we introduce a new class of structured matrices, which unifies and generalizes structured classes from previous works. We examine properties of this class and build a structured orthogonal parametrization upon it. We then use this parametrization to modify the orthogonal fine-tuning framework, improving parameter and computational efficiency. We empirically validate our method on different domains, including adapting of text-to-image diffusion models and downstream task fine-tuning in language modeling. Additionally, we adapt our construction for orthogonal convolutions and conduct experiments with 1-Lipschitz neural networks.

📄 PDF Abstract BibTeX arXiv:2406.10019

Code (1)

skonor/group_and_shuffle pytorch

Tasks

Computational EfficiencyLanguage ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OrthoFuse: Training-free Riemannian Fusion of Orthogonal Style-Concept Adapters for Diffusion Models

2026-04-06 · Ali Aliev, Kamil Garifullin, Nikolay Yudin, Vera Soboleva 외 arxiv

In a rapidly growing field of model training there is a constant practical interest in parameter-efficient fine-tuning and various techniques that use a small amount of training data to adapt the model to a narrow task. …

parameter-efficient fine-tuning

Cheap Orthogonal Constraints in Neural Networks: A Simple Parametrization of the Orthogonal and Unitary Group

2019-01-24 · Mario Lezcano-Casado, David Martínez-Rubio

We introduce a novel approach to perform first-order optimization with orthogonal and unitary constraints. This approach is based on a parametrization stemming from Lie group theory through the exponential map. The param…

Dynamic Shuffle: An Efficient Channel Mixture Method

2023-10-04 · Kaijun Gong, Zhuowen Yin, Yushu Li, Kailing Guo 외

The redundancy of Convolutional neural networks not only depends on weights but also depends on inputs. Shuffling is an efficient operation for mixing channel information but the shuffle order is usually pre-defined. To …

Binarizationimage-classificationImage Classification

CWY Parametrization: a Solution for Parallelized Optimization of Orthogonal and Stiefel Matrices

2020-04-18 · Valerii Likhosherstov, Jared Davis, Krzysztof Choromanski, Adrian Weller

We introduce an efficient approach for optimization over orthogonal groups on highly parallel computation units such as GPUs or TPUs. As in earlier work, we parametrize an orthogonal matrix as a product of Householder re…

Machine TranslationTranslationVideo Prediction

Structured Sparsification with Joint Optimization of Group Convolution and Channel Shuffle

2020-02-19 · Xin-Yu Zhang, Kai Zhao, Taihong Xiao, Ming-Ming Cheng 외

Recent advances in convolutional neural networks(CNNs) usually come with the expense of excessive computational overhead and memory footprint. Network compression aims to alleviate this issue by training compact models w…

Network Pruning