Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information
Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalLanguage ModellingSimilar Papers 제목 키워드 기반
Why Chinese Web-as-Corpus is Wacky? Or: How Big Data is Killing Chinese Corpus Linguistics
This paper aims to examine and evaluate the current development of using Web-as-Corpus (WaC) paradigm in Chinese corpus linguistics. I will argue that the unstable notion of wordhood in Chinese and the resulting diverse …
Chinese Word SegmentationChinese Preposition Selection for Grammatical Error Diagnosis
Misuse of Chinese prepositions is one of common word usage errors in grammatical error diagnosis. In this paper, we adopt the Chinese Gigaword corpus and HSK corpus as L1 and L2 corpora, respectively. We explore gated re…
Language ModelingLanguage ModellingSentenceTAGLSICC: A Large Scale Informal Chinese Corpus
Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a l…
Chinese Word SegmentationDeep LearningSentiment AnalysisCSL: A Large-scale Chinese Scientific Literature Dataset
Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development …
text-classificationText ClassificationHyperbolic Deep Learning for Chinese Natural Language Understanding
Recently hyperbolic geometry has proven to be effective in building embeddings that encode hierarchical and entailment information. This makes it particularly suited to modelling the complex asymmetrical relationships be…
Chinese Word SegmentationDeep LearningNatural Language Understanding