paper-with-me

홈 › Papers

When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it

2022-05-06 · NAACL 2022 7 · Sebastian Schuster, Tal Linzen

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases (NPs), such as 'a dog', frequently introduce discourse entities but this behavior is modulated by sentential operators such as negation. For example, 'a dog' in 'Arthur doesn't own a dog' does not introduce a discourse entity due to the presence of negation. In this work, we adapt the psycholinguistic assessment of language models paradigm to higher-level linguistic phenomena and introduce an English evaluation suite that targets the knowledge of the interactions between sentential operators and indefinite NPs. We use this evaluation suite for a fine-grained investigation of the entity tracking abilities of the Transformer-based models GPT-2 and GPT-3. We find that while the models are to a certain extent sensitive to the interactions we investigate, they are all challenged by the presence of multiple NPs and their behavior is not systematic, which suggests that even models at the scale of GPT-3 do not fully acquire basic entity tracking abilities.

📄 PDF Abstract BibTeX arXiv:2205.03472

Code (1)

sebschu/discourse-entity-lm 공식 구현

Tasks

NegationSentence

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Discriminative Fine-Tuning Discriminative Fine-Tuning is a fine-tuning strategy that is used for ULMFiT type models. Instead of using the same learning rate…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

When a sentence does not introduce a discourse entity, Transformer-based models still often refer to it

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases, such as 'a dog', frequently introduce discourse entities but th…

NegationSentence

Entity-Augmented Distributional Semantics for Discourse Relations

2014-12-17 · Yangfeng Ji, Jacob Eisenstein

Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A mo…

RelationSentence

An entity-driven recursive neural network model for chinese discourse coherence modeling

2017-04-14 · Fan Xu, Shujing Du, Maoxi Li, Mingwen Wang

Chinese discourse coherence modeling remains a challenge taskin Natural Language Processing field.Existing approaches mostlyfocus on the need for feature engineering, whichadoptthe sophisticated features to capture the l…

Coherence EvaluationFeature EngineeringMachine TranslationSentence+2

Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing

2026-06-25 · Mreedul Gupta, Advait Deshmukh, Ashwin Umadi, Matt Pauk 외 arxiv

Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail. We hypothesize that a key limitation is the reliance on sentence-level contex…

Entity Typing

New or Old? Exploring How Pre-Trained Language Models Represent Discourse Entities

2022-10-01 · COLING 2022 10 · Sharid Loáiciga, Anne Beyer, David Schlangen

Recent research shows that pre-trained language models, built to generate text conditioned on some context, learn to encode syntactic knowledge to a certain degree. This has motivated researchers to move beyond the sente…

Binary ClassificationSentence