paper-with-me

홈 › Papers

Temporal Attention for Language Models

2022-02-04 · Findings (NAACL) 2022 7 · Guy D. Rosin, Kira Radinsky

Pretrained language models based on the transformer architecture have shown great success in NLP. Textual training data often comes from the web and is thus tagged with time-specific information, but most language models ignore this information. They are trained on the textual data alone, limiting their ability to generalize temporally. In this work, we extend the key component of the transformer architecture, i.e., the self-attention mechanism, and propose temporal attention - a time-aware self-attention mechanism. Temporal attention can be applied to any transformer model and requires the input texts to be accompanied with their relevant time points. It allows the transformer to capture this temporal information and create time-specific contextualized word representations. We leverage these representations for the task of semantic change detection; we apply our proposed mechanism to BERT and experiment on three datasets in different languages (English, German, and Latin) that also vary in time, size, and genre. Our proposed model achieves state-of-the-art results on all the datasets.

📄 PDF Abstract BibTeX arXiv:2202.02093

Code (1)

guyrosin/temporal_attention 공식 구현 pytorch

Tasks

Change Detection

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Weight Decay 설명 없음
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Adam 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…

Similar Papers 제목 키워드 기반

SLGTformer: An Attention-Based Approach to Sign Language Recognition

2022-12-21 · Neil Song, Yu Xiang

Sign language is the preferred method of communication of deaf or mute people, but similar to any language, it is difficult to learn and represents a significant barrier for those who are hard of hearing or unable to spe…

Sign Language Recognition

STARK: Spatio-Temporal Attention for Representation of Keypoints for Continuous Sign Language Recognition

2026-03-17 · Suvajit Patra, Soumitra Samanta arxiv

Continuous Sign Language Recognition (CSLR) is a crucial task for understanding the languages of deaf communities. Contemporary keypoint-based approaches typically rely on spatio-temporal encoding, where spatial interact…

Sign Language Recognition

Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability

2025-10-09 · Chengzhi Li, Heyan Huang, Ping Jian, Zhen Yang 외 arxiv

Large language models (LLMs) often generate self-contradictory outputs, which severely impacts their reliability and hinders their adoption in practical applications. In video-language models (Video-LLMs), this phenomeno…

Spatio-Temporal Ranked-Attention Networks for Video Captioning

2020-01-17 · Anoop Cherian, Jue Wang, Chiori Hori, Tim K. Marks

Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features…

Video Captioning

Temporal Attention Modules for Memory-Augmented Neural Networks

2021-01-01 · Rodolfo Palma, Alvaro Soto, Luis Martí, Nayat Sanchez-pi

We introduce two temporal attention modules which can be plugged into traditional memory augmented recurrent neural networks to improve their performance in natural language processing tasks. The temporal attention modul…