paper-with-me

홈 › Papers

Effective Use of Transformer Networks for Entity Tracking

2019-09-05 · IJCNLP 2019 11 · Aditya Gupta, Greg Durrett

Tracking entities in procedural language requires understanding the transformations arising from actions on entities as well as those entities' interactions. While self-attention-based pre-trained language encoders like GPT and BERT have been successfully applied across a range of natural language understanding tasks, their ability to handle the nuances of procedural texts is still untested. In this paper, we explore the use of pre-trained transformer networks for entity tracking tasks in procedural text. First, we test standard lightweight approaches for prediction with pre-trained transformers, and find that these approaches underperform even simple baselines. We show that much stronger results can be attained by restructuring the input to guide the transformer model to focus on a particular entity. Second, we assess the degree to which transformer networks capture the process dynamics, investigating such factors as merged entities and oblique entity references. On two different tasks, ingredient detection in recipes and QA over scientific processes, we achieve state-of-the-art results, but our models still largely attend to shallow context clues and do not form complex representations of intermediate entity or process state.

📄 PDF Abstract BibTeX arXiv:1909.02635

Code (1)

aditya2211/transformer-entity-tracking 공식 구현 pytorch

Tasks

Natural Language Understanding

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

Chain and Causal Attention for Efficient Entity Tracking

2024-10-07 · Erwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre Allauzen

This paper investigates the limitations of transformers for entity-tracking tasks in large language models. We identify a theoretical constraint, showing that transformers require at least $\log_2 (n+1)$ layers to handle…

Language ModelingLanguage Modelling

When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it

2022-05-06 · NAACL 2022 7 · Sebastian Schuster, Tal Linzen

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases (NPs), such as 'a dog', frequently introduce discourse entities …

NegationSentence

When a sentence does not introduce a discourse entity, Transformer-based models still often refer to it

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases, such as 'a dog', frequently introduce discourse entities but th…

NegationSentence

TrackFormer: Multi-Object Tracking with Transformers

2021-01-07 · CVPR 2022 1 · Tim Meinhardt, Alexander Kirillov, Laura Leal-Taixe, Christoph Feichtenhofer

The challenging task of multi-object tracking (MOT) requires simultaneous reasoning about track initialization, identity, and spatio-temporal trajectories. We formulate this task as a frame-to-frame set prediction proble…

DecoderMulti-Object TrackingObjectObject Tracking+1

Model Optimization for Multi-Camera 3D Detection and Tracking

2026-01-31 · Ethan Anderson, Justin Silva, Kyle Zheng, Sameer Pusegaonkar 외 arxiv

Outside-in multi-camera perception is increasingly important in indoor environments, where networks of static cameras must support multi-target tracking under occlusion and heterogeneous viewpoints. We evaluate Sparse4D,…