paper-with-me

Papers

Refining Targeted Syntactic Evaluation of Language Models

2021-04-19 · NAACL 2021 4 · Benjamin Newman, Kai-Siang Ang, Julia Gong, John Hewitt

Targeted syntactic evaluation of subject-verb number agreement in English (TSE) evaluates language models' syntactic knowledge using hand-crafted minimal pairs of sentences that differ only in the main verb's conjugation. The method evaluates whether language models rate each grammatical sentence as more likely than its ungrammatical counterpart. We identify two distinct goals for TSE. First, evaluating the systematicity of a language model's syntactic knowledge: given a sentence, can it conjugate arbitrary verbs correctly? Second, evaluating a model's likely behavior: given a sentence, does the model concentrate its probability mass on correctly conjugated verbs, even if only on a subset of the possible verbs? We argue that current implementations of TSE do not directly capture either of these goals, and propose new metrics to capture each goal separately. Under our metrics, we find that TSE overestimates systematicity of language models, but that models score up to 40% better on verbs that they predict are likely in context.

📄 PDF Abstract BibTeX arXiv:2104.09635

Code (1)

bnewm0609/refining-tse 공식 구현 pytorch

Tasks

Sentence

Similar Papers 제목 키워드 기반

SyntaxGym: An Online Platform for Targeted Evaluation of Language Models

2020-07-01 · ACL 2020 6 · Jon Gauthier, Jennifer Hu, Ethan Wilcox, Peng Qian 외

Targeted syntactic evaluations have yielded insights into the generalizations learned by neural network language models. However, this line of research requires an uncommon confluence of skills: both the theoretical know…

Experimental DesignLanguage ModelingLanguage Modelling

Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

2024-11-12 · Daria Kryvosheieva, Roger Levy

Language models (LMs) are capable of acquiring elements of human-like syntactic knowledge. Targeted syntactic evaluation tests have been employed to measure how well they form generalizations about syntactic phenomena in…

Diversity

Language model acceptability judgements are not always robust to context

2022-12-18 · Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra 외

Targeted syntactic evaluations of language models ask whether models show stable preferences for syntactically acceptable content over minimal-pair unacceptable inputs. Most targeted syntactic evaluation datasets ask mod…

In-Context LearningLanguage ModelingLanguage ModellingSentence

Structured Sentiment Analysis as Dependency Graph Parsing

2021-05-30 · ACL 2021 5 · Jeremy Barnes, Robin Kurtz, Stephan Oepen, Lilja Øvrelid 외

Structured sentiment analysis attempts to extract full opinion tuples from a text, but over time this task has been subdivided into smaller and smaller sub-tasks, e,g,, target extraction or targeted polarity classificati…

Sentiment Analysis

Targeted Paraphrasing on Deep Syntactic Layer for MT Evaluation

2015-08-01 · WS 2015 8 · Petra Baran{\v{c}}{\'\i}kov{\'a}, Rudolf Rosa
Machine Translation