paper-with-me

Papers

Language model acceptability judgements are not always robust to context

2022-12-18 · Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams

Targeted syntactic evaluations of language models ask whether models show stable preferences for syntactically acceptable content over minimal-pair unacceptable inputs. Most targeted syntactic evaluation datasets ask models to make these judgements with just a single context-free sentence as input. This does not match language models' training regime, in which input sentences are always highly contextualized by the surrounding corpus. This mismatch raises an important question: how robust are models' syntactic judgements in different contexts? In this paper, we investigate the stability of language models' performance on targeted syntactic evaluations as we vary properties of the input context: the length of the context, the types of syntactic phenomena it contains, and whether or not there are violations of grammaticality. We find that model judgements are generally robust when placed in randomly sampled linguistic contexts. However, they are substantially unstable for contexts containing syntactic structures matching those in the critical test content. Among all tested models (GPT-2 and five variants of OPT), we significantly improve models' judgements by providing contexts with matching syntactic structures, and conversely significantly worsen them using unacceptable contexts with matching but violated syntactic structures. This effect is amplified by the length of the context, except for unrelated inputs. We show that these changes in model performance are not explainable by simple features matching the context and the test inputs, such as lexical overlap and dependency overlap. This sensitivity to highly specific syntactic features of the context can only be explained by the models' implicit in-context learning abilities.

📄 PDF Abstract BibTeX arXiv:2212.08979

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningLanguage ModelingLanguage ModellingSentence

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

The Influence of Context on Sentence Acceptability Judgements

2018-07-01 · ACL 2018 7 · Jean-Philippe Bernardy, Shalom Lappin, Jey Han Lau

We investigate the influence that document context exerts on human acceptability judgements for English sentences, via two sets of experiments. The first compares ratings for sentences presented on their own with ratings…

Language ModelingLanguage ModellingMachine TranslationSentence+1

Work Hard, Play Hard: Collecting Acceptability Annotations through a 3D Game

2022-06-01 · LREC 2022 6 · Federico Bonetti, Elisa Leonardelli, Daniela Trotta, Raffaele Guarasci 외

Corpus-based studies on acceptability judgements have always stimulated the interest of researchers, both in theoretical and computational fields. Some approaches focused on spontaneous judgements collected through diffe…

CoLA

Human acceptability judgements for extractive sentence compression

2019-02-01 · Abram Handler, Brian Dillon, Brendan O'Connor

Recent approaches to English-language sentence compression rely on parallel corpora consisting of sentence-compression pairs. However, a sentence may be shortened in many different ways, which each might be suited to the…

SentenceSentence Compression

Revisiting Acceptability Judgements

2023-05-23 · Hai Hu, Ziyin Zhang, Weifang Huang, Jackie Yan-Ki Lai 외

In this work, we revisit linguistic acceptability in the context of large language models. We introduce CoLAC - Corpus of Linguistic Acceptability in Chinese, the first large-scale acceptability dataset for a non-Indo-Eu…

Cross-Lingual TransferLinguistic Acceptability

The Acceptability Delta Criterion: Testing Knowledge of Language using the Gradience of Sentence Acceptability

2021-11-01 · EMNLP (BlackboxNLP) 2021 11 · Héctor Vázquez Martínez

Any test that promises to assess Human Knowledge of Language (KoL) for any statistically-based Language Model (LM) must meet three requirements: (1) comprehensive coverage of linguistic phenomena; (2) replicable and stat…

Language ModellingSentence