paper-with-me

홈 › Papers

Sparsity Emerges Naturally in Neural Language Models

2019-07-22 · ICML Workshop Deep_Phenomen 2019 6 · Naomi Saphra, Adam Lopez

Concerns about interpretability, computational resources, and principled inductive priors have motivated efforts to engineer sparse neural models for NLP tasks. If sparsity is important for NLP, might well-trained neural models naturally become roughly sparse? Using the Taxi-Euclidean norm to measure sparsity, we find that frequent input words are associated with concentrated or sparse activations, while frequent target words are associated with dispersed activations but concentrated gradients. We find that gradients associated with function words are more concentrated than the gradients of content words, even controlling for word frequency.

📄 PDF Abstract BibTeX arXiv:1908.01817

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Attention is Naturally Sparse with Gaussian Distributed Input

2024-04-03 · Yichuan Deng, Zhao Song, Chiwun Yang

The computational intensity of Large Language Models (LLMs) is a critical bottleneck, primarily due to the $O(n^2)$ complexity of the attention mechanism in transformer architectures. Addressing this, sparse attention em…

Computational Efficiency

Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models

2025-07-23 · Andrii Balashov arxiv

Reinforcement learning (RL) is a key post-pretraining step for aligning large language models (LLMs) with complex tasks and human preferences. While it is often assumed that RL fine-tuning requires updating most of a mod…

Reinforcement Learning

Sparse Attention as Compact Kernel Regression

2026-01-30 · Saul Santos, Nuno Gonçalves, Daniel C. McNamee, Marcos Treviso 외 arxiv

Recent work has revealed a link between self-attention mechanisms in transformers and test-time kernel regression via the Nadaraya-Watson estimator, with standard softmax attention corresponding to a Gaussian kernel. How…

Density Estimation

Towards the Connection between Activation Sparsity and Flat Minima

2026-05-25 · Ze Peng, Jian Zhang, Lei Qi, Yang Gao 외 arxiv

The observation that activation sparsity emerges in MLP blocks of standardly trained Transformers offers an opportunity to drastically reduce computation costs without sacrificing performance. To theoretically explain th…

Sparse Attention with Linear Units

2021-04-14 · EMNLP 2021 11 · Biao Zhang, Ivan Titov, Rico Sennrich

Recently, it has been argued that encoder-decoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants. In this work, we introduce a novel, simple method for a…

DecoderDiversityMachine TranslationTranslation+1