paper-with-me

홈 › Papers

Do Long-Range Language Models Actually Use Long-Range Context?

2021-09-19 · EMNLP 2021 11 · Simeng Sun, Kalpesh Krishna, Andrew Mattarella-Micke, Mohit Iyyer

Language models are generally trained on short, truncated input sequences, which limits their ability to use discourse-level information present in long-range context to improve their predictions. Recent efforts to improve the efficiency of self-attention have led to a proliferation of long-range Transformer language models, which can process much longer sequences than models of the past. However, the ways in which such models take advantage of the long-range context remain unclear. In this paper, we perform a fine-grained analysis of two long-range Transformer language models (including the \emph{Routing Transformer}, which achieves state-of-the-art perplexity on the PG-19 long-sequence LM benchmark dataset) that accept input sequences of up to 8K tokens. Our results reveal that providing long-range context (i.e., beyond the previous 2K tokens) to these models only improves their predictions on a small set of tokens (e.g., those that can be copied from the distant context) and does not help at all for sentence-level prediction tasks. Finally, we discover that PG-19 contains a variety of different document types and domains, and that long-range context helps most for literary novels (as opposed to textbooks or magazines).

📄 PDF Abstract BibTeX arXiv:2109.09115

Code (0)

등록된 구현이 없습니다.

Tasks

2k8kSentence

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Addressing Deep Learning Model Uncertainty in Long-Range Climate Forecasting with Late Fusion

2021-12-10 · Ken C. L. Wong, Hongzhi Wang, Etienne E. Vos, Bianca Zadrozny 외

Global warming leads to the increase in frequency and intensity of climate extremes that cause tremendous loss of lives and property. Accurate long-range climate prediction allows more time for preparation and disaster r…

Management

How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections

2022-06-24 · Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra 외

Linear time-invariant state space models (SSM) are a classical model from engineering and statistics, that have recently been shown to be very promising in machine learning through the Structured State Space sequence mod…

Long-range modelingState Space Models

Long-Range Correlation Underlying Childhood Language and Generative Models

2017-12-11 · Kumiko Tanaka-Ishii

Long-range correlation, a property of time series exhibiting long-term memory, is mainly studied in the statistical physics domain and has been reported to exist in natural language. Using a state-of-the-art method for s…

Time SeriesTime Series Analysis

How Does CP Length Affect the Sensing Range for OFDM-ISAC?

2025-03-11 · Xiaoli Xu, Zhiwen Zhou, Yong Zeng

Orthogonal frequency division multiplexing (OFDM), which has been the dominating waveform for contemporary wireless communications, is also regarded as a competitive candidate for future integrated sensing and communicat…

Integrated sensing and communicationISAC

Long-Range Biometric Identification in Real World Scenarios: A Comprehensive Evaluation Framework Based on Missions

2024-09-03 · Deniz Aykac, Joel Brogan, Nell Barber, Ryan Shivers 외

The considerable body of data available for evaluating biometric recognition systems in Research and Development (R\&D) environments has contributed to the increasingly common problem of target performance mismatch. Biom…

Face Recognition