paper-with-me

홈 › Papers

Tensor Train Low-rank Approximation (TT-LoRA): Democratizing AI with Accelerated LLMs

2024-08-02 · Afia Anjum, Maksim E. Eren, Ismael Boureima, Boian Alexandrov, Manish Bhattarai

In recent years, Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing (NLP) tasks, such as question-answering, sentiment analysis, text summarization, and machine translation. However, the ever-growing complexity of LLMs demands immense computational resources, hindering the broader research and application of these models. To address this, various parameter-efficient fine-tuning strategies, such as Low-Rank Approximation (LoRA) and Adapters, have been developed. Despite their potential, these methods often face limitations in compressibility. Specifically, LoRA struggles to scale effectively with the increasing number of trainable parameters in modern large scale LLMs. Additionally, Low-Rank Economic Tensor-Train Adaptation (LoRETTA), which utilizes tensor train decomposition, has not yet achieved the level of compression necessary for fine-tuning very large scale models with limited resources. This paper introduces Tensor Train Low-Rank Approximation (TT-LoRA), a novel parameter-efficient fine-tuning (PEFT) approach that extends LoRETTA with optimized tensor train (TT) decomposition integration. By eliminating Adapters and traditional LoRA-based structures, TT-LoRA achieves greater model compression without compromising downstream task performance, along with reduced inference latency and computational overhead. We conduct an exhaustive parameter search to establish benchmarks that highlight the trade-off between model compression and performance. Our results demonstrate significant compression of LLMs while maintaining comparable performance to larger models, facilitating their deployment on resource-constraint platforms.

📄 PDF Abstract BibTeX arXiv:2408.01008

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationModel Compressionparameter-efficient fine-tuningQuestion AnsweringSentiment AnalysisText Summarization

Similar Papers 제목 키워드 기반

Tensor-Efficient High-Dimensional Q-learning

2025-11-05 · Junyi Wu, Dan Li arxiv

High-dimensional reinforcement learning(RL) faces challenges with complex calculations and low sample efficiency in large state-action spaces. Q-learning algorithms struggle particularly with the curse of dimensionality,…

Reinforcement Learning

Quantum-Enhanced LLM Efficient Fine Tuning

2025-03-17 · Xiaofei Kong, Lei LI, Menghan Dou, Zhaoyun Chen 외

Low-Rank Adaptation (LoRA) enables efficient fine-tuning of pre-trained language models via low-rank matrix approximation, which is effective in many scenarios. However, its low-rank representation capacity is constraine…

parameter-efficient fine-tuning

Fast Tucker Rank Reduction for Non-Negative Tensors Using Mean-Field Approximation

2021-03-04 · NeurIPS 2021 12 · Kazu Ghalamkari, Mahito Sugiyama

We present an efficient low-rank approximation algorithm for non-negative tensors. The algorithm is derived from our two findings: First, we show that rank-1 approximation for tensors can be viewed as a mean-field approx…

Tensor Decomposition

Mode-wise Tensor Decompositions: Multi-dimensional Generalizations of CUR Decompositions

2021-03-19 · HanQin Cai, Keaton Hamm, Longxiu Huang, Deanna Needell

Low rank tensor approximation is a fundamental tool in modern machine learning and data science. In this paper, we study the characterization, perturbation analysis, and an efficient sampling strategy for two primary ten…

Near-Linear Time and Fixed-Parameter Tractable Algorithms for Tensor Decompositions

2022-07-15 · Arvind V. Mahankali, David P. Woodruff, Ziyu Zhang

We study low rank approximation of tensors, focusing on the tensor train and Tucker decompositions, as well as approximations with tree tensor networks and more general tensor networks. For tensor train decomposition, we…

Dimensionality ReductionTensor DecompositionTensor Networks