paper-with-me

Papers

Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

2012-05-01 · LREC 2012 5 · Chi-Hsin Yu, Yi-jie Tang, Hsin-Hsi Chen

Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalLanguage Modelling

Similar Papers 제목 키워드 기반

Why Chinese Web-as-Corpus is Wacky? Or: How Big Data is Killing Chinese Corpus Linguistics

2014-05-01 · LREC 2014 5 · Shu-Kai Hsieh

This paper aims to examine and evaluate the current development of using Web-as-Corpus (WaC) paradigm in Chinese corpus linguistics. I will argue that the unstable notion of wordhood in Chinese and the resulting diverse …

Chinese Word Segmentation

Chinese Preposition Selection for Grammatical Error Diagnosis

2016-12-01 · COLING 2016 12 · Hen-Hsen Huang, Yen-Chi Shao, Hsin-Hsi Chen

Misuse of Chinese prepositions is one of common word usage errors in grammatical error diagnosis. In this paper, we adopt the Chinese Gigaword corpus and HSK corpus as L1 and L2 corpora, respectively. We explore gated re…

Language ModelingLanguage ModellingSentenceTAG

LSICC: A Large Scale Informal Chinese Corpus

2018-11-26 · Jianyu Zhao, Zhuoran Ji

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a l…

Chinese Word SegmentationDeep LearningSentiment Analysis

CSL: A Large-scale Chinese Scientific Literature Dataset

2022-09-12 · COLING 2022 10 · Yudong Li, Yuqing Zhang, Zhe Zhao, Linlin Shen 외

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development …

text-classificationText Classification

Hyperbolic Deep Learning for Chinese Natural Language Understanding

2018-12-11 · Marko Valentin Micic, Hugo Chu

Recently hyperbolic geometry has proven to be effective in building embeddings that encode hierarchical and entailment information. This makes it particularly suited to modelling the complex asymmetrical relationships be…

Chinese Word SegmentationDeep LearningNatural Language Understanding