paper-with-me

홈 › Papers

CLUECorpus2020: A Large-scale Chinese Corpus for Pre-training Language Model

2020-03-03 · Liang Xu, Xuanwei Zhang, Qianqian Dong

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G raw corpus with 35 billion Chinese characters, which is retrieved from Common Crawl. To better understand this corpus, we conduct language understanding experiments on both small and large scale, and results show that the models trained on this corpus can achieve excellent performance on Chinese. We release a new Chinese vocabulary with a size of 8K, which is only one-third of the vocabulary size used in Chinese Bert released by Google. It saves computational cost and memory while works as good as original vocabulary. We also release both large and tiny versions of the pre-trained model on this corpus. The former achieves the state-of-the-art result, and the latter retains most precision while accelerating training and prediction speed for eight times compared to Bert-base. To facilitate future work on self-supervised learning on Chinese, we release our dataset, new vocabulary, codes, and pre-trained models on Github.

📄 PDF Abstract BibTeX arXiv:2003.01355

Code (2)

CLUEbenchmark/CLUECorpus2020 공식 구현 tf
CLUEbenchmark/CLUEPretrainedModels tf

Tasks

8kLanguage ModelingLanguage ModellingSelf-Supervised LearningText Generation

Methods 이 논문이 사용한 방법론

Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Residual Connection 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Weight Decay 설명 없음
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…

Similar Papers 제목 키워드 기반

LSICC: A Large Scale Informal Chinese Corpus

2018-11-26 · Jianyu Zhao, Zhuoran Ji

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a l…

Chinese Word SegmentationDeep LearningSentiment Analysis

Development of a Web-Scale Chinese Word N-gram Corpus with Parts of Speech Information

2012-05-01 · LREC 2012 5 · Chi-Hsin Yu, Yi-jie Tang, Hsin-Hsi Chen

Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variat…

Information RetrievalLanguage Modelling

BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark

2023-02-18 · Dakuan Lu, Hengkui Wu, Jiaqing Liang, Yipei Xu 외

To advance Chinese financial natural language processing (NLP), we introduce BBT-FinT5, a new Chinese financial pre-training language model based on the T5 model. To support this effort, we have built BBT-FinCorpus, a la…

Language ModelingLanguage Modelling

Ancient-Modern Chinese Translation with a Large Training Dataset

2018-08-11 · Dayiheng Liu, Jiancheng Lv, Kexin Yang, Qian Qu

Ancient Chinese brings the wisdom and spirit culture of the Chinese nation. Automatic translation from ancient Chinese to modern Chinese helps to inherit and carry forward the quintessence of the ancients. However, the l…

Cultural Vocal Bursts Intensity PredictionMachine TranslationNMTTranslation

Incorporating translation quality estimation into Chinese-Korean neural machine translation

2021-08-01 · CCL 2021 8 · Li Feiyu, Zhao Yahui, Yang Feiyang, Cui Rongyi

“Exposure bias and poor translation diversity are two common problems in neural machine trans-lation (NMT) which are caused by the general of the teacher forcing strategy for training inthe NMT models. Moreover the NMT m…

Machine TranslationNMTTranslation