Ellipsis in Chinese AMR Corpus
Ellipsis is very common in language. It{'}s necessary for natural language processing to restore the elided elements in a sentence. However, there{'}s only a few corpora annotating the ellipsis, which draws back the automatic detection and recovery of the ellipsis. This paper introduces the annotation of ellipsis in Chinese sentences, using a novel graph-based representation Abstract Meaning Representation (AMR), which has a good mechanism to restore the elided elements manually. We annotate 5,000 sentences selected from Chinese TreeBank (CTB). We find that 54.98{\%} of sentences have ellipses. 92{\%} of the ellipses are restored by copying the antecedents{'} concepts. and 12.9{\%} of them are the new added concepts. In addition, we find that the elided element is a word or phrase in most cases, but sometimes only the head of a phrase or parts of a phrase, which is rather hard for the automatic recovery of ellipsis.
Code (0)
등록된 구현이 없습니다.
Tasks
Abstract Meaning RepresentationSentenceSimilar Papers 제목 키워드 기반
NoEl: An Annotated Corpus for Noun Ellipsis in English
Ellipsis resolution has been identified as an important step to improve the accuracy of mainstream Natural Language Processing (NLP) tasks such as information retrieval, event extraction, dialog systems, etc. Previous co…
Event ExtractionInformation RetrievalPOSRetrievalTo Drop or Not to Drop? Predicting Argument Ellipsis Judgments: A Case Study in Japanese
Speakers sometimes omit certain arguments of a predicate in a sentence; such omission is especially frequent in pro-drop languages. This study addresses a question about ellipsis -- what can explain the native speakers' …
Language ModelingLanguage ModellingSentenceTowards Handling Verb Phrase Ellipsis in English-Hindi Machine Translation
English-Hindi machine translation systems have difficulty interpreting verb phrase ellipsis (VPE) in English, and commit errors in translating sentences with VPE. We present a solution and theoretical backing for the tre…
Machine TranslationSentenceTranslationRiSAWOZ: A Large-Scale Multi-Domain Wizard-of-Oz Dataset with Rich Semantic Annotations for Task-Oriented Dialogue Modeling
In order to alleviate the shortage of multi-domain data and to capture discourse phenomena for task-oriented dialogue modeling, we propose RiSAWOZ, a large-scale multi-domain Chinese Wizard-of-Oz dataset with Rich Semant…
Dialogue State TrackingIntent DetectionNatural Language Understandingslot-filling+2