paper-with-me

홈 › Papers

TRAMS: Training-free Memory Selection for Long-range Language Modeling

2023-10-24 · Haofei Yu, Cunxiang Wang, Yue Zhang, Wei Bi

The Transformer architecture is crucial for numerous AI models, but it still faces challenges in long-range language modeling. Though several specific transformer architectures have been designed to tackle issues of long-range dependencies, existing methods like Transformer-XL are plagued by a high percentage of ineffective memories. In this study, we present a plug-and-play strategy, known as TRAining-free Memory Selection (TRAMS), that selects tokens participating in attention calculation based on one simple metric. This strategy allows us to keep tokens that are likely to have a high attention score with the current queries and ignore the other ones. We have tested our approach on the word-level benchmark (WikiText-103) and the character-level benchmark (enwik8), and the results indicate an improvement without having additional training or adding additional parameters.

📄 PDF Abstract BibTeX arXiv:2310.15494

Code (1)

lwaekfjlk/trams 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Multi-Head Attention 설명 없음
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Adaptive Input Representations Adaptive Input Embeddings extend the adaptive softmax to input word representations. The factorization assigns more…
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs

2026-07-16 · Shahrzad Esmat, Dhawal Shah, Ali Jannesari arxiv

The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) scor…

Deep and interpretable regression models for ordinal outcomes

2020-10-16 · Lucas Kook, Lisa Herzog, Torsten Hothorn, Oliver Dürr 외

Outcomes with a natural order commonly occur in prediction tasks and often the available input data are a mixture of complex data like images and tabular predictors. Deep Learning (DL) models are state-of-the-art for ima…

image-classificationImage Classificationregression

OmniMem: Perturbation-aware Memory Compression for Streaming Audio-Visual LLMs

2026-05-26 · Guangzhi Sun, Yixuan Li, Yudong Yang, Chao Zhang arxiv

Audio-visual large language models (LLMs) hold strong promise for long-form video understanding, yet their long-video inference is fundamentally limited by the linear growth of video tokens and key-value (KV) caches. We …

AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations

2026-03-02 · Cheng Jiayang, Dongyu Ru, Lin Qiu, Yiyang Li 외 arxiv

Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory. Existing memory benchmarks rely on st…

SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree

2024-10-21 · Shuangrui Ding, Rui Qian, Xiaoyi Dong, Pan Zhang 외

The Segment Anything Model 2 (SAM 2) has emerged as a powerful foundation model for object segmentation in both images and videos, paving the way for various downstream video applications. The crucial design of SAM 2 for…

Heuristic SearchObjectSegmentationSemantic Segmentation+4