A Japanese Masked Language Model for Academic Domain
We release a pretrained Japanese masked language model for an academic domain. Pretrained masked language models have recently improved the performance of various natural language processing applications. In domains such as medical and academic, which include a lot of technical terms, domain-specific pretraining is effective. While domain-specific masked language models for medical and SNS domains are widely used in Japanese, along with domain-independent ones, pretrained models specific to the academic domain are not publicly available. In this study, we pretrained a RoBERTa-based Japanese masked language model on paper abstracts from the academic database CiNii Articles. Experimental results on Japanese text classification in the academic domain revealed the effectiveness of the proposed model over existing pretrained models.
Code (1)
Tasks
ArticlesLanguage ModelingLanguage Modellingmodeltext-classificationText ClassificationSimilar Papers 제목 키워드 기반
Domain-Specific Japanese ELECTRA Model Using a Small Corpus
Recently, domain shift, which affects accuracy due to differences in data between source and target domains, has become a serious issue when using machine learning methods to solve natural language processing tasks. With…
ArticlesComputational EfficiencyDocument ClassificationLanguage Modeling+2Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs
Why do we build local large language models (LLMs)? What should a local LLM learn from the target language? Which abilities can be transferred from other languages? Do language-specific scaling laws exist? To explore the…
Arithmetic ReasoningCode GenerationQuestion AnsweringReading ComprehensionAutomatic Assistance for Academic Word Usage
This paper describes a writing assistance system that helps students improve their academic writing. Given an input text, the system suggests lexical substitutions that aim to incorporate more academic vocabulary. The su…
Language ModelingLanguage ModellingPatton: Language Model Pretraining on Text-Rich Networks
A real-world text corpus sometimes comprises not only text documents but also semantic links between them (e.g., academic papers in a bibliographic network are linked by citations and co-authorships). Text documents and …
Language ModelingLanguage ModellingMasked Language Modelingmodel+1UnihanLM: Coarse-to-Fine Chinese-Japanese Language Model Pretraining with the Unihan Database
Chinese and Japanese share many characters with similar surface morphology. To better utilize the shared knowledge across the languages, we propose UnihanLM, a self-supervised Chinese-Japanese pretrained masked language …
Language ModelingLanguage Modelling