paper-with-me

홈 › Papers

NOPE: A Corpus of Naturally-Occurring Presuppositions in English

2021-09-14 · CoNLL (EMNLP) 2021 11 · Alicia Parrish, Sebastian Schuster, Alex Warstadt, Omar Agha, Soo-Hwan Lee, Zhuoye Zhao, Samuel R. Bowman, Tal Linzen

Understanding language requires grasping not only the overtly stated content, but also making inferences about things that were left unsaid. These inferences include presuppositions, a phenomenon by which a listener learns about new information through reasoning about what a speaker takes as given. Presuppositions require complex understanding of the lexical and syntactic properties that trigger them as well as the broader conversational context. In this work, we introduce the Naturally-Occurring Presuppositions in English (NOPE) Corpus to investigate the context-sensitivity of 10 different types of presupposition triggers and to evaluate machine learning models' ability to predict human inferences. We find that most of the triggers we investigate exhibit moderate variability. We further find that transformer-based models draw correct inferences in simple cases involving presuppositions, but they fail to capture the minority of exceptional cases in which human judgments reveal complex interactions between context and triggers.

📄 PDF Abstract BibTeX arXiv:2109.06987

Code (1)

nyu-mll/nope 공식 구현

Similar Papers 제목 키워드 기반

Automatic Extraction of Clausal Embedding Based on Large-Scale English Text Data

2025-06-16 · Iona Carslaw, Sivan Milton, Nicolas Navarre, Ciyang Qing 외

For linguists, embedded clauses have been of special interest because of their intricate distribution of syntactic and semantic features. Yet, current research relies on schematically created language examples to investi…

Constituency Parsing

Learning Syntax from Naturally-Occurring Bracketings

2021-04-28 · NAACL 2021 4 · Tianze Shi, Ozan İrsoy, Igor Malioutov, Lillian Lee

Naturally-occurring bracketings, such as answer fragments to natural language questions and hyperlinks on webpages, can reflect human syntactic intuition regarding phrasal boundaries. Their availability and approximate c…

Constituency Parsing

Collecting Natural SMS and Chat Conversations in Multiple Languages: The BOLT Phase 2 Corpus

2014-05-01 · LREC 2014 5 · Zhiyi Song, Stephanie Strassel, Haejoong Lee, Kevin Walker 외

The DARPA BOLT Program develops systems capable of allowing English speakers to retrieve and understand information from informal foreign language genres. Phase 2 of the program required large volumes of naturally occurr…

Machine TranslationTranslation

VietMix: A Naturally Occurring Vietnamese-English Code-Mixed Corpus with Iterative Augmentation for Machine Translation

2025-05-30 · Hieu Tran, Phuong-Anh Nguyen-Le, Huy Nghiem, Quang-Nhan Nguyen 외

Machine translation systems fail when processing code-mixed inputs for low-resource languages. We address this challenge by curating VietMix, a parallel corpus of naturally occurring code-mixed Vietnamese text paired wit…

Machine TranslationSynthetic Data GenerationTranslation

CREPE: Open-Domain Question Answering with False Presuppositions

2022-11-30 · Xinyan Velocity Yu, Sewon Min, Luke Zettlemoyer, Hannaneh Hajishirzi

Information seeking users often pose questions with false presuppositions, especially when asking about unfamiliar topics. Most existing question answering (QA) datasets, in contrast, assume all questions have well defin…

Open-Domain Question AnsweringQuestion Answering