The RST Spanish-Chinese Treebank
Discourse analysis is necessary for different tasks of Natural Language Processing (NLP). As two of the most spoken languages in the world, discourse analysis between Spanish and Chinese is important for NLP research. This paper aims to present the first open Spanish-Chinese parallel corpus annotated with discourse information, whose theoretical framework is based on the Rhetorical Structure Theory (RST). We have evaluated and harmonized each annotation part to obtain a high annotated-quality corpus. The corpus is already available to the public.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
SCTB: A Chinese Treebank in Scientific Domain
Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in thi…
Chinese Word SegmentationMachine TranslationTranslationThe IULA Spanish LSP Treebank
This paper presents the IULA Spanish LSP Treebank, a dependency treebank of over 41,000 sentences of different domains (Law, Economy, Computing Science, Environment, and Medicine), developed in the framework of the Europ…
The CUHK Discourse TreeBank for Chinese: Annotating Explicit Discourse Connectives for the Chinese TreeBank
The lack of open discourse corpus for Chinese brings limitations for many natural language processing tasks. In this work, we present the first open discourse treebank for Chinese, namely, the Discourse Treebank for Chin…
Part-Of-Speech TaggingQuestion AnsweringSentenceSentence Compression+2How does discourse affect Spanish-Chinese Translation? A case study based on a Spanish-Chinese parallel corpus
With their huge speaking populations in the world, Spanish and Chinese occupy important positions in linguistic studies. Since the two languages come from different language systems, the translation between Spanish and C…
TranslationA Dependency Treebank of the Chinese Buddhist Canon
We present a dependency treebank of the Chinese Buddhist Canon, which contains 1,514 texts with about 50 million Chinese characters. The treebank was created by an automatic parser trained on a smaller treebank, containi…
Dependency ParsingPart-Of-Speech Tagging