paper-with-me

홈 › Papers

NoEl: An Annotated Corpus for Noun Ellipsis in English

2020-05-01 · LREC 2020 5 · Payal Khullar, Kushal Majmundar, Manish Shrivastava

Ellipsis resolution has been identified as an important step to improve the accuracy of mainstream Natural Language Processing (NLP) tasks such as information retrieval, event extraction, dialog systems, etc. Previous computational work on ellipsis resolution has focused on one type of ellipsis, namely Verb Phrase Ellipsis (VPE) and a few other related phenomenon. We extend the study of ellipsis by presenting the No(oun)El(lipsis) corpus - an annotated corpus for noun ellipsis and closely related phenomenon using the first hundred movies of Cornell Movie Dialogs Dataset. The annotations are carried out in a standoff annotation scheme that encodes the position of the licensor, the antecedent boundary, and Part-of-Speech (POS) tags of the licensor and antecedent modifier. Our corpus has 946 instances of exophoric and endophoric noun ellipsis, making it the biggest resource of noun ellipsis in English, to the best of our knowledge. We present a statistical study of our corpus with novel insights on the distribution of noun ellipsis, its licensors and antecedents. Finally, we perform the tasks of detection and resolution of noun ellipsis with different classifiers trained on our corpus and report baseline results.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Event ExtractionInformation RetrievalPOSRetrieval

Similar Papers 제목 키워드 기반

Announcing Prague Czech-English Dependency Treebank 2.0

2012-05-01 · LREC 2012 5 · Jan Haji{\v{c}}, Eva Haji{\v{c}}ov{\'a}, Jarmila Panevov{\'a}, Petr Sgall 외

We introduce a substantial update of the Prague Czech-English Dependency Treebank, a parallel corpus manually annotated at the deep syntactic layer of linguistic representation. The English part consists of the Wall Stre…

Coreference ResolutionSentence

Exploring Statistical and Neural Models for Noun Ellipsis Detection and Resolution in English

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Payal Khullar

Computational approaches to noun ellipsis resolution has been sparse, with only a naive rule-based approach that uses syntactic feature constraints for marking noun ellipsis licensors and selecting their antecedents. In …

Towards Handling Verb Phrase Ellipsis in English-Hindi Machine Translation

2019-12-01 · ICON 2019 12 · Niyati Bafna, Dipti Sharma

English-Hindi machine translation systems have difficulty interpreting verb phrase ellipsis (VPE) in English, and commit errors in translating sentences with VPE. We present a solution and theoretical backing for the tre…

Machine TranslationSentenceTranslation

DELA Corpus - A Document-Level Corpus Annotated with Context-Related Issues

2021-11-01 · WMT (EMNLP) 2021 11 · Sheila Castilho, João Lucas Cavalheiro Camargo, Miguel Menezes, Andy Way

Recently, the Machine Translation (MT) community has become more interested in document-level evaluation especially in light of reactions to claims of “human parity”, since examining the quality at the level of the docum…

Machine TranslationSentenceTranslation

ParCor 1.0: A Parallel Pronoun-Coreference Corpus to Support Statistical MT

2014-05-01 · LREC 2014 5 · Liane Guillou, Christian Hardmeier, Aaron Smith, J{\"o}rg Tiedemann 외

We present ParCor, a parallel corpus of texts in which pronoun coreference ― reduced coreference in which pronouns are used as referring expressions ― has been annotated. The corpus is intended to be used both as a resou…

Machine TranslationTranslation